SetupScore research

You are reading the scores wrong

We all read an 85 as high. Nearly every score lands between 80 and 92, which puts 85 near the middle of the pile.

By Carlos VitorinoUpdated 7 min readReview score analysis

A dot distribution drawn on the full 0 to 100 scale. Each dot is one of the 99 products in the set. Almost all of them stack into a narrow cluster between 80 and 92 near the right hand end, and the rest of the scale is empty. A dot distribution drawn on the full 0 to 100 scale. Each dot is one of the 99 products in the set. Almost all of them stack into a narrow cluster between 80 and 92 near the right hand end, and the rest of the scale is empty.

You see an 85 on a monitor and file it away as good. Not the best, but clear of the middle, with room above and plenty below. That habit comes from school, and reading a review score like a school grade is the most common mistake a shopper makes.

The mistake is not that 85 is a low number. It is that the scale you are picturing, the one running from 0 to 100 with failure at the bottom, does not exist in practice. SetupScore's set holds 99 products, each with at least 2 verified expert ratings. Nothing in it scored below 54 or above 93.

Why do most review scores land in the same narrow band?

Half of that scale is empty, and the half that remains is narrower than it looks. Most of the field crowds into a single run of scores, from 80 to 92, which means more than four products in five sit in a twelve point band. SetupScore's analysis of 99 product scores found 82.8% of them packed into a twelve point stretch of the scale. Only 16 products, 16.2% of the set, sit below 80.

There is not much room left to separate 99 products. The median score is 85, the exact number most shoppers treat as respectable. Above it sit 45 products and below it 44, with 10 more landing on 85 itself. A tenth of the field is tied on a single value, which is what a scale does once it runs out of room, and it means an 85 is not clear of the middle. It is the middle.

Almost the whole axis is empty

Figure 1. All 99 products plotted across the full 0 to 100 axis. The shaded band, 80 to 92, holds 82.8% of them, and 45 of them sit in the five points between 85 and 89. Nothing in the set sits below 54, so most of the axis is decoration.

Tight packing changes what a small difference is worth, and near the middle 2 points can move you past more than a fifth of the field. Across the 99 products SetupScore has scored, an 85 beats 44.4% of the field. An 87 beats 65.7% of them. That gap is worth 21.3 points of rank. The same 2 points near the top would barely move a product.

The top of the range has almost nothing left to move through. Exactly 1 product in the whole set scores above 92, and a score of 92 already beats 98.0% of everything we cover. The bottom edge is just as thin, where a score of 75 beats 7.1% of the field. Almost everything else sits in the middle.

A score is a position, not a grade

score

75 7.1%
80 16.2%
82 28.3%
85 44.4%
87 65.7%
90 89.9%
92 98.0%
95 100%

share of the field this score beats

View the numbers as a table
Review scores converted into percentile rank across SetupScore's product set.
scoreshare of the field this score beats
757.1%
8016.2%
8228.3%
8544.4%
8765.7%
9089.9%
9298.0%
95100%
Figure 2. The same scores converted into rank. Read across from a score to find the share of the field it beats. An 85 beats 44.4% of the field. An 87, only 2 points higher, beats 65.7%. The middle of the scale is steep, and those 2 points are worth 21.3 points of rank.

Two forces produce a field this tight, and the effect has a name: score compression. The first is selection, or range restriction: the products that get reviewed are already survivors, picked because readers are already shopping for them, which strips out the broken and the obscure before scoring starts. A field of finalists does not spread out.

The second force works on the numbers themselves. Reviewers publish on different scales, and a five point scale and a hundred point scale do not mean the same thing. One offers a handful of steps, the other a hundred.

A scale with 5 steps has only 5 available answers, and converting it to a 100 point axis does not create more. A rating of 4 out of 5 becomes an 80. A rating of 4.5 out of 5 becomes a 90. Coarse scales can only land on a handful of fixed values, and those values sit near the middle of the axis.

There is an obvious objection here: we may have created the clustering ourselves. We publish an average of several ratings per product, and averaging always pulls numbers closer together. Two things answer that objection.

The first is that the clustering was there before we averaged anything. According to SetupScore's analysis of 611 expert ratings, 65.0% of them land inside a ten point range. The individual ratings behave the same way on their own, and 397 of them fall between 80 and 90. The middle half spans the same 10 points. Those raw ratings run the full width of the axis, from 10 to 100, while the finished scores only cover 54 to 93. Averaging narrows the picture. It does not create the crowding inside it.

The second is a test of that worry. We dealt the 611 real ratings out at random into imaginary products of the same sizes, and repeated that 4,000 times. Each time we measured how far apart the middle half of those imaginary products landed, the gap between the 25th and the 75th percentile of their rating averages. If averaging were doing the work, the imaginary products would sit as far apart as the real ones. They did not. The random deals came out tighter, an interquartile range of 5.35 points on average against 6.85 for the real products, and only 34 of the 4,000 reassignments reached the real figure (p = 0.0087, seed 20260811). Averaging does squeeze the numbers, but it is not what put these products so close together.

So the number is not a grade. It is a position in a short, tightly packed queue, and it behaves like one. An 85 does not mean 85 percent of anything, and it does not mean the product got 85 percent of the way to some ideal. It means several people who tested it landed near the middle of a narrow band, and none of them found a reason to push it higher.

Once you see the queue, the number changes job. A score is a rank, not a measurement. It does not tell you how good a product is in some absolute sense. It tells you where that product stands in a line of 99.

Read the distance, not the number

A position in that line is only useful if you know what it is worth. Figure 2 turns each score into a percentile rank, the share of the field it beats, so you can read across and see where any score really sits. The shape matters more than any single row: the steps are crowded through the middle of the scale and stretched almost flat at the top.

The practical move is to compare two scores against each other rather than one score against an imaginary 100. A gap of 3 points between two products near the middle is a large difference in rank. The same 3 points above 92 covers almost nothing, because almost nothing is there.

Money does not follow the top of the scale in the set we cover. Taken within category, never pooled, since different price floors let a change in the mix imitate a price effect, the link is undetectable: monitors r = -0.034 (p = 0.831) across 41 priced products, keyboards r = 0.139 (p = 0.463) across 30, headphones r = 0.233 (p = 0.252) across 26. None is significant, so price does not predict rank among the products we have scored. The medians are flat or lower in the higher band each time: monitors $849.50 to $700, keyboards $135 to $119.50, headphones $299 either side.

What a higher score costs, within a category

category and score band

Monitors, below 85 $849.50
Monitors, 85 and above $700
Keyboards, below 85 $135
Keyboards, 85 and above $119.50
Headphones, below 85 $299
Headphones, 85 and above $299

median price within category, US dollars

Figure 3. Median price below 85 and at 85 and above, within each category, for the 98 priced products we have scored. The median is flat or lower in the higher band every time: monitors $849.50 to $700, keyboards $135 to $119.50, headphones $299 to $299. None of the three within-category correlations is significant, so this set shows no detectable relationship rather than evidence that cheaper products score better.

We are paid a percentage of the sale price, so we have an incentive for expensive products to score well. The per-category figures are the test of that, and across these 98 priced products the test does not show it happening: price and score move independently, and the higher band is no more expensive than the lower one. With 26 to 41 products per category, that is an absence of a detectable effect here, not proof that none exists.

How should you read a review score?

None of this makes scores useless. It makes them a narrower instrument than they look, and there are better questions to ask of one. Start with position. Out of everything that reviewer has rated, where does this product sit? Any outlet with a back catalogue has built a line of its own, and the score in front of you has a place in it. That place is the information. The distance to 100 is not.

How the number was reached changes what it is worth, so it pays to work out where the score came from. Some are the output of a system, fed by measurements and repeatable across products. Others are a verdict, reached in one move by somebody who lived with the thing for a week. Both are worth reading. They are not the same claim.

A decimal is a cheap tell. Seeing 8.7 rather than 9 probably means arithmetic happened somewhere, though it settles nothing on its own, because a finer opinion is still an opinion. What it does suggest is that the number was built out of parts.

An overall score is a blend, and the weights inside it belong to the reviewer. A pair of headphones can lose ground on portability because the cups do not fold. That is a real fault for somebody carrying them onto a train every morning, and no fault at all for somebody who listens at a desk. The total hides the difference. Aspect scores show it, which is why the parts are worth more to you than the total.

How much evidence sits behind a score matters as much as the score itself. A number resting on 2 opinions is not the same object as one resting on 16, even when both land on an 85. Our own set thins out fast at that end. Of the 99 products SetupScore covers, 5 rest on the minimum of 2 ratings, and the typical product here is carried by only 6. That is a thin base by any standard, and on those 5 a single extra review can move the score a long way. Ask that of any score you meet, including ours.

How we handle it here

This is the problem SetupScore was built around. Scores from different outlets do not share an axis, so they are normalized onto one common scale before anything is combined, which is why a 4 out of 5 and an 80 can land in the same column. Most of what feeds a page carries no score at all. Only 611 of the 2,395 sources behind these 99 products arrive with a number attached, and the rest, largely YouTube and Reddit, are read as text and tone. Each product is then split into named aspects with their own scores and quotes, so the weighting can be yours instead of ours. None of that removes the need to read. It only makes the number easier to argue with.