Most people now decide what to watch partly on an aggregate score. The scores measure different things, none measures quality directly, and the differences matter.
The two main approaches
Worth separating because they are frequently confused.
One approach converts each review into a binary judgement — positive or negative — and reports the percentage that were positive. This is a measure of consensus, not of enthusiasm.
The other assigns each review a numerical score and reports a weighted average. This measures average assessed quality and is closer to what people assume they are reading.
The distinction produces very different results for the same film. A work everybody finds moderately good scores extremely well under the first system and moderately under the second.
A divisive work — half the critics consider it remarkable, half consider it a failure — scores poorly under the first and moderately under the second, and the score conveys nothing about the division.
What gets lost
The specific information that would actually help.
Variance. Whether opinions clustered or diverged is frequently the most useful thing to know, and no single number carries it.
Direction of disagreement. A film that critics dislike and audiences love, or the reverse, is telling you something specific about who it is for.
Intensity. A review that finds a film competent and one that finds it extraordinary can register identically.
And the reasons, which are the entire content of the reviews being aggregated.
The critic and audience gap
Worth interpreting rather than treating as a scandal.
Large divergences between critical and audience scores are common and are frequently informative.
Critics see far more films, which makes them more sensitive to originality and less tolerant of familiarity. Audiences self-select into films they expect to enjoy, which biases their scores upward.
Genre matters. Horror and comedy consistently score better with audiences than with critics; certain kinds of drama do the reverse.
Which means the gap is a signal about what kind of film it is rather than evidence that one party is wrong.
The manipulation problem
Real and specific to audience scores.
Review bombing — coordinated negative rating, frequently before release and for reasons unrelated to the work — has affected a number of titles.
Platforms have responded with verification requirements, delayed opening of audience scoring, and weighting toward confirmed viewers.
Those measures help and do not eliminate it, and an audience score in the first days after a contentious release should be treated with caution.
The commercial pressure on the number
An underdiscussed issue.
Aggregate scores affect box office measurably, which means studios have a substantial interest in them.
Which has produced questions about which outlets are included in aggregation, how borderline reviews are classified, and the influence of access relationships on the reviews being aggregated in the first place.
None of this suggests fabrication. It does mean the number sits inside a commercial system rather than outside it.
What I actually use instead
Having largely stopped reading the headline figure.
Two or three critics whose taste I have calibrated against my own over years, which is worth more than any aggregate because I know how to correct for them.
The distribution rather than the average, where a platform shows it. A histogram tells you immediately whether a work is divisive.
Negative reviews of things I am inclined to like, which are more informative than positive ones because they identify the specific failure modes.
And the descriptive portion of any review rather than the verdict, since knowing what a thing is like is more useful than knowing whether somebody approved.
The one legitimate use
Aggregates are genuinely good at identifying the extremes.
Something scoring very low across many reviews is reliably poor. Something scoring very high across many is reliably at least competent.
The vast middle, which is where most films are, is where the number stops carrying information, and that is exactly where people rely on it most.
The sample problem
A limitation worth understanding.
Aggregates for widely released films draw on many reviews. Aggregates for smaller releases frequently draw on very few, and a score derived from eight reviews is not comparable to one derived from three hundred.
Most platforms indicate the count and few readers look at it.
The same applies to audience scores, where a small number of highly motivated raters can move a figure substantially, particularly early.
Checking the count takes a second and changes how much weight the number deserves more than any other single consideration.
Rating your own
A small practice that has been more useful than reading anybody else's scores.
Recording a short note after watching something — a sentence about what it was and whether it worked — builds a personal record that is calibrated to your own taste by definition.
After a couple of years it becomes genuinely useful for spotting patterns in what you actually enjoy, which frequently differ from what you believe you enjoy.
Mine revealed that I rate ambitious failures higher than competent successes, which explains a number of recommendations that have not gone well.