Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Interesting. To summarize their methodology:

- Take the top 20 hotels from trip advisor

- Filter out all non-five star reviews, plus any non-english, excessively short, or first-time reviews. Then sample (via a log normal distribution on review length) 20 reviews per hotel. This is their "real" dataset.

- Use Mechanical Turk to collect 400 reviews for these hotels. Turkers are instructed to pretend they work for the marketing dept of the hotel and are to write deceptively fake reviews. Again, quality filters on length, user approval rating, and deduplication are applied. Turkers are paid $1 per review. This creates their "fake" dataset.

I suppose one could still argue that there are selection bias issues here. The sample size is also moderate. Nevertheless, it's a novel approach and you have to start somewhere. Interesting work.



It seems like writing a naive Bayesean classifier for "Was this written by a Turker" should be like taking candy from a baby who hates candy, has very slippery fingers, and is unconscious.


Till the person who hires the Turkers includes said filter along with the work order.

"Write a good that passes this, this and this filter by a fair margin"

You might try seeing who's hiring Turkers for what. It might give you an idea how much filtering is needed.


If it was that easy, why did their human judges fail at it?


Does the experimental setup for the human judges sound fair to you?

For example, the naive Bayes classifier knows the a priori distribution of review spam (which appears to be held to 50%), but do the undergraduate human judges? It would appear not, given that one judge only labeled 12% deceptive.

Likewise, were the human judges able to see examples of truthful and deceptive reviews before beginning the task? (In other words, are the human judges solving a different problem, e.g., "deception detection", than the classifier e.g., "similarity to prior deceptive reviews from Turkers").

If these are differences between the human and computer annotator setups, are they major differences? Can you spot any other big differences between the two experimental setups?


Yeah, there are some selection bias issues here. From what I have seen, anomaly detection especially graph based ones can be pretty robust, but they break when the person has written only one or so reviews- which is what I guess the majority of reviews are from.


What's the point of filtering out all non-five star reviews? I for one would be very interested in fake reviews by my competitors.


They mention negative deceptive review detection as further work.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: