Hamid Reza Mohagheghi Contact Hamid

Online Reviews and the Limits of Social Information

Review scores are treated as measurements. They are better understood as the output of a social process with known biases.

Star ratings are treated as measurements of quality. They are better understood as the output of a social process — one with documented selection effects, herding dynamics and strategic manipulation. The information is real. So is the distortion, and the distortion is systematic rather than random.

Reviews carry genuine economic weight

Chevalier and Mayzlin (2006) compared book sales across two retailers and found that an improvement in reviews at one site increased relative sales there, establishing a causal effect rather than a correlation. Luca (2016) found that a one-star increase in a restaurant's Yelp rating produced a measurable revenue increase, concentrated among independent businesses rather than chains.

The chain finding is instructive. Where a consumer already has strong prior information about quality, reviews add little. Their value is highest under uncertainty, which is also where their distortions do the most damage.

Who writes reviews is not who buys

Review populations are self-selected, and the selection is bimodal: people with unusually good and unusually bad experiences are most motivated to report. The resulting distribution has a well-known J shape that no amount of averaging corrects, because the missing middle was never sampled.

Verified-purchase filters help with fraud but not with this. A verified reviewer is still a self-selected one. Any programme that increases review volume by soliciting responses changes the composition of the sample, which is why average ratings often shift when a solicitation programme launches — the product did not change.

Herding

Ratings are not independent observations. Reviewers see prior ratings before contributing, and prior ratings influence what they write. Muchnik, Aral and Taylor (2013) showed experimentally that a single artificially added positive vote produced a lasting increase in final scores, while a negative one was corrected.

The consequence is that a rating aggregates a sequence, not a sample. Early reviews carry disproportionate influence over the final figure, and early reviews are the ones most likely to come from atypical customers or from the operator's own network.

Manipulation is measurable

Mayzlin, Dover and Chevalier (2014) compared hotel reviews across platforms with different verification requirements and found patterns of promotional and defamatory reviewing consistent with strategic manipulation, concentrated where it was cheapest to do and where the incentive was strongest — independent hotels with a neighbouring competitor.

The rate of outright fake review is contested and platform-dependent. What is not contested is that the incentive exists and that detection is imperfect.

Reading reviews well

The useful information in a review corpus is generally not the average. It is the distribution shape, the recency profile, and the content of the specific complaints. Three detailed one-star reviews describing the same failure mode carry more information than the difference between 4.3 and 4.5 stars, which is within the noise of the sampling process.

For businesses, the corresponding discipline is to read reviews as a diagnostic rather than a scoreboard. The score is partly an artefact. The recurring specific complaint is data.

What this means in practice

Treat aggregate ratings as a weak signal with known biases, weight recent and specific content heavily, and be aware that any intervention which changes who reviews will change the score without changing the product. The general dynamics behind these effects are covered in social proof.

References

  1. Chevalier, J. A., & Mayzlin, D. (2006). The effect of word of mouth on sales: Online book reviews. Journal of Marketing Research, 43(3), 345–354.
  2. Luca, M. (2016). Reviews, reputation, and revenue: The case of Yelp.com. Harvard Business School Working Paper 12-016.
  3. Mayzlin, D., Dover, Y., & Chevalier, J. (2014). Promotional reviews: An empirical investigation of online review manipulation. American Economic Review, 104(8), 2421–2455.
  4. Muchnik, L., Aral, S., & Taylor, S. J. (2013). Social influence bias: A randomized experiment. Science, 341(6146), 647–651.
  5. Hu, N., Pavlou, P. A., & Zhang, J. (2009). Overcoming the J-shaped distribution of product reviews. Communications of the ACM, 52(10), 144–147.