Review aggregation was a genuinely good idea. Instead of relying on one critic whose taste might not match yours, look at many and find the consensus. Reduce it to a number for convenience.
The convenience turned out to have costs that are larger than they look, and the way these scores now function in the industry is meaningfully different from what they measure.
The binary conversion
The central flaw in the best-known model. Each review is classified as positive or negative, and the score is the percentage positive.
Which means a review saying "flawed but fascinating, worth seeing" and one saying "an unqualified masterpiece" count identically. So do "mildly disappointing" and "an insult to the audience".
The consequence is that the score measures consensus rather than quality. A film that everybody agrees is perfectly fine scores higher than one that half the critics consider a masterpiece and half consider a failure.
That's a defensible thing to measure. It is not what audiences think they're reading, and the difference matters enormously for exactly the kind of ambitious, divisive work that most deserves attention.
The weighted-average alternative
The other main model assigns each review a numerical score and averages them, weighted by outlet.
Better in principle, and it has its own problems. Many critics don't give scores, so the aggregator assigns one based on the text — a judgement call, sometimes disputed by the critics themselves.
And the weighting is opaque. Which outlets count for more, and by how much, is not fully disclosed, which makes the number difficult to interpret.
Which reviews get counted
An under-examined issue. Aggregators decide which outlets are eligible, and that decision shapes the result.
The criteria have shifted over time and have been criticised from several directions — for excluding smaller or newer publications, and separately for including outlets whose reviewing is thin. Either way, the composition of the sample is a set of editorial decisions that determine the number, and most people reading a score have no idea a sample selection occurred.
There's also a timing effect. Early reviews come disproportionately from outlets attending festivals or press events, which skews the initial score. A film can open at a high number and settle considerably lower as wider reviews arrive.
What it did to the industry
This is where it stops being an academic complaint.
Scores now materially affect commercial performance. Studios track them closely, and there's evidence that they influence opening weekend figures, particularly for films without strong brand recognition.
The predictable responses have followed. Careful management of review embargoes so scores appear at chosen moments. Selective early access for outlets expected to be favourable. And a general awareness in marketing departments that a number crossing a threshold is worth real money.
More concerning is the effect on what gets made. If a divisive film scores worse than a bland one, and the score affects performance, the incentive points towards blandness. Nobody decides that explicitly, and the pressure is real.
The audience score problem
A separate mess. Audience scores are vulnerable to coordinated manipulation, and there have been documented cases of review-bombing — organised campaigns to depress a score for reasons unrelated to the work.
Aggregators have introduced verification measures to address this, with mixed results. The underlying problem is that anyone can leave a rating for something they haven't seen, and any system permitting that will be gamed.
The gap between critic and audience scores gets treated as meaningful — evidence of critics being out of touch, or audiences having poor taste, depending on your priors. Often it's just evidence that one number is a sample of professionals and the other is a sample of whoever showed up.
How to actually use them
My suggestions, having thought about this more than is healthy.
Read the distribution rather than the average, where it's available. A film with reviews clustered around "good" is a different proposition from one with reviews at both extremes, and the second is usually more interesting.
Find three or four critics whose taste you've calibrated against your own and weight them heavily. This is what the whole aggregation project was meant to replace and it remains more useful.
Treat scores between roughly sixty and eighty percent as containing almost no information. That's where most things land and the differences within that band are noise.
And be especially sceptical of a very high score. It frequently indicates something inoffensive that nobody disliked, which is not the same as something you'll be glad you watched.
What critics themselves say
It is worth noting that a substantial number of working critics dislike this system, and several have written about it at length. The complaints are consistent: a nuanced piece reduced to a binary classification, a considered reservation flattened into a thumbs-down, and the resulting number then used to characterise their opinion in ways they would not recognise.
Some publications have responded by declining to supply scores at all. That helps the individual critic and does very little to the aggregate, since the aggregator can infer a classification from the text regardless.
Which is the underlying issue. A system that converts writing into numbers without the writer's consent is going to produce numbers the writer disagrees with, and there is no mechanism by which they can object.