RankquantRQ

Methodology

HotelDanAriMelJenTimAvg
A5/55/53/53/54.0
B5/55/55/53/53/54.2
C4/54/54.0
D4/55/51/52/53.0
IrrelevantMost important information

Summary

Old rating system (Amazon, Booking.com, TripAdvisor):
raw score average
Hotel B4.2
Hotel A4.0
Hotel C4.0
Hotel D3.0
Rankquant
Each reviewer is re-centred on their own average, then scored as a percentile.
Hotel C92(z +1.42)
Hotel A51(z +0.02)
Hotel B49(z −0.02)
Hotel D8(z −1.41)

Follow Hotel B: ringed at 4.2, first on the raw average — and underlined at 49, third once every reviewer is re-centred.

Other sites would have you choose Hotel B. Rankquant clearly shows you math-perfected Hotel C is superior.

Who is doing this, and why the math

Rankquant is a team of statistically-trained editors. We are not a review site that hired a data scientist; we are applied statisticians who believe the way review scores are published today is straightforwardly broken. Our job on this page is to explain, in plain English paired with the exact equations, how we turn messy real-world reviews into a number you can actually trust.

None of this math is novel. Per-reviewer normalization, reviewer fixed-effects aggregation, and shrinking a thin sample toward the population mean are textbook techniques in psychometrics, meta-analysis, and sports analytics — see /theory/founding-metrics for a first-principles tour of the five primitives this pipeline is built from. What's different is publishing them, applied rigorously, as the basis of a consumer review site.

The problem every review site has

Every major review surface suffers severe right-skew inflation. Raw averages cannot distinguish a genuinely exceptional product from a merely average one.

4.4/5 avg

Amazon's long-run average across billions of reviews across all categories.

Marketplace Pulse review analyses

4.2/5 avg

Yelp's long-run category-averaged restaurant rating.

Yelp transparency reports

8.4/10 avg

Booking.com's long-run hotel rating; hotels below 8.0 are flagged as low-rated.

Booking.com scoring bands

When everything is "excellent," the word carries no information. And there's a second problem hiding underneath the inflation: different reviewers use the same numbers to mean different things. One reviewer's 88 is another reviewer's ceiling; one's 3-stars is another's "I liked it." Averaging raw numbers across reviewers mixes those private scales together and throws away signal.

Our fix has two moves. First, normalize every reviewer onto their own z-scale so personal grading habits wash out. Second, count every one of those reviewers equally and let the size of the sample — not our opinion of the reviewer — decide how much confidence the result earns.

The method, in three steps

Step 1 — per-reviewer z-score normalization

For each reviewer u, compute the reviewer's personal mean μu and standard deviation σu over all of their reviews across all products in our database. Their rating of product i is then converted to a z-score:

z_{u,i}  =  ( r_{u,i}  −  μ_u )  /  σ_u

  where
    r_{u,i} = reviewer u's raw rating of product i
    μ_u     = reviewer u's personal mean rating across all their reviews
    σ_u     = reviewer u's personal standard deviation (Bessel-corrected, df = n_u − 1)
    n_u     = number of reviews reviewer u has written in our dataset

Two reviewers who disagree on scale but agree on quality produce identical z-scores. A reviewer who always rates 90–96 and a reviewer who always rates 75–88 will both produce z ≈ +1.3for their personal favorites — because that's where each reviewer sits relative to their own distribution. The z-score is dimensionless. It's the unit reviewers actually share.

Who counts as a "qualifying reviewer"?

Not every reviewer's opinions belong in the z-scale. Two conditions:

A reviewer with nu ≥ 2 but σu = 0isn't useless — they just can't be normalized the standard way. We keep them on file, and in the verticals that admit them (Products does — see Scope) we impute a plausible σ from their source's pooled dispersion. When they are admitted they enter the same unweighted mean as everyone else, at exactly the same weight.

Cross-category pooling. A reviewer who writes both wine and bourbon reviews goes into one reviewer pool with one μu and one σucomputed across everything they've rated. That's deliberate: a person's personal scale is a property of the person, not the product category. Pooling keeps nu large and σu stable.

Step 2 — combine the reviewers, weighting none of them

After Step 1, every item has a list of z-scores — one per qualifying reviewer. We take their plain arithmetic mean. That is the whole of Step 2.

mean_z(i)  =  (1 / N_i)  ·  Σ_{u ∈ Q_i}  z_{u,i}       ← equal weight per reviewer

  Q_i = qualifying reviewers of item i  (n_u ≥ 2, σ_u > 0)
  N_i = |Q_i|

Every qualifying reviewer counts the same.An amateur with five reviews counts exactly as much as a professional with five thousand. There is no credibility weight, no source weight, no per-reviewer multiplier and no editorial thumb anywhere on the scale. The only thing that changes a reviewer's influence is how many other reviewers are in the average with them.

This is a deliberate constraint rather than a missing feature. The moment a weight exists, the interesting question stops being "what did the reviewers say" and becomes "who decided the weights, and what would change if they were different" — and that question is unanswerable from outside. A single unweighted mean is the version of this number a reader can reconstruct from published data without trusting our judgment about anybody.

We do not sell weight, and we do not assign it either. Every qualifying reviewer counts once.

Editorial standard

Earlier versions of this page described a second, source-weighted aggregate and a third "broadened" one, presented as three lenses. Those were designed but never shipped: no published score has ever depended on a source weight, and there has never been a live weight table. The description has been removed rather than left standing as a claim about a thing that does not exist. What Step 3 does publish is three rankings of this one number — not three different numbers.

Step 3 — rank that one number, three ways

Step 2 produced a single quantity per item. Step 3 does not produce more quantities — it ranks that one quantity three times, and those three rankings are the three percentiles on every detail page.

The three published percentiles. All three are rankings of the same unweighted mean z-score — none of them weights a reviewer.
Global percentileThe empirical-CDF rank of the mean z-score against every item we rank in that vertical, expressed 0–100. This is the headline. It is built from the mean z-score and nothing else — no confidence floor, no shrinkage. It reports the measurement as made.
In-cohort percentileThe same mean z-score, re-ranked inside a narrower peer group — same category and a close price band, or city and nightly-rate band for hotels. Same quantity, smaller field. A re-ranking, not a second computation.
AI-adjusted percentileThe same mean z-score after it is discounted for how much confidence the sample supports: multiplied by n/(n+53), so thin samples are pulled toward the corpus average. One line of arithmetic, applied identically in all six verticals — no model, no training, no inference. This is the only one of the three that adjusts anything.
The three published percentiles. All three are rankings of the same unweighted mean z-score — none of them weights a reviewer.

Charging a score for its uncertainty

This is the part that matters most for what you see on a page — and it happens in exactly one of the three figures. An item with 4 reviewers averaging z = +2.1 might be exceptional, or its mean might be noise; one with 80 reviewers averaging z = +1.6 has paid its statistical dues. The global and in-cohort percentiles do not arbitrate that — they rank the mean as measured, and +2.1 goes ahead of +1.6. The arbitration lives in the AI-adjusted percentile, which discounts each item in proportion to how little evidence stands behind it. Note where the penalty lands: on the sample, never on the reviewer. We never decide that a particular person's opinion is worth less; we decide how much confidence a given number of opinions buys.

The discount is one multiplication — mean_z · n/(n+53) — applied identically in all six verticals, and it is written out in full under the AI-adjusted percentile below. Separately, and ordering nothing, we publish the lower bound of a 90% one-tailed confidence intervalaround each item's mean, wherever the sample defines one, as a diagnostic. That floor answers a question no percentile does: "given this sample size, what is a defensibly pessimistic estimate of the item's true quality?"

Ẑ       = mean_z(i), the unweighted mean from Step 2
SE(Ẑ)   = 1 / √N                                 (standard error of the mean)
floor   = Ẑ  −  1.645 · SE(Ẑ)                    ← one-tailed 90% CI lower bound

  N = the number of qualifying reviewers in the mean. Because every reviewer
      carries the same weight, the effective sample size IS N — there is no
      weighted correction to apply.

  Published per item — ciFloorZ in the catalog exports, zCiLower and zMargin
  on products — and read next to the score. Undefined at N = 1, where it is
  absent rather than faked. No percentile on the site is ranked by it.

Worked example — two real-world shapes, and what each figure does with them:

ProductN (reviewers)Mean Ẑ — what global ranksn/(n+53)Adjusted z — what AI-adjusted ranks90% floor (published, ranks nothing)
Thin-sample darling4+2.100.070+0.15+1.28
Well-reviewed consensus80+1.600.602+0.96+1.42

On the mean — which is what the global and in-cohort percentiles rank — the darling wins, +2.10 to +1.60, and we publish that rather than quietly correcting it. On the AI-adjusted percentile the order reverses, +0.96 to +0.15, because four reviewers keep only 7% of their measured distance from average while eighty keep 60%. The published confidence floors, +1.28 and +1.42, point the same way without ordering anything. This is by design, and it's the same intuition Bayesian sports ratings, IMDb's Top 250 formula, and Wilson score confidence intervals (used by Reddit and Yelp internally) all use: penalize uncertainty, reward consistency — kept in its own figure instead of folded silently into the headline.

From a z-score to a 0–100 percentile

A mean z-score is fine for math but means nothing to a reader. So we rank every item's value against every other item's and express the result as a percentile using the empirical CDF:

p_global(i)  =  100 · rank( value(i) )  /  N_total

  value(i) = mean_z(i)          in all six verticals — the measurement itself,
                                with no confidence floor and no shrinkage in it
  rank()   = ordinal rank, ties split at midpoint
  N_total  = number of items with a valid value in that vertical

  An item with no mean_z is absent from the denominator and stays unranked. It
  is never parked at 0 or at the corpus average to fill the column.

A percentile of 90 means the item scores higher than 90% of everything else we rank in that vertical. 50 is the median. 10 means it sits in the bottom 10%.

Cohort percentiles — a re-ranking, not a re-computation

Global percentile answers "how does this product compare to everything we measure?" — which is useful but sometimes unfair. A $14 bottle that beats all other $14 bottles is doing exactly what a $14 bottle should do; showing it at the 25th global percentile (against $200 Burgundy) hides that achievement.

So we also publish a cohort percentile. It is the same value, re-ranked within a narrower peer group.

Cohort(i) = { j : category(j) = category(i)
                AND |price(j) − price(i)| / price(i) ≤ 0.20 }

p_cohort(i) = 100 · rank( value(i) ) / |Cohort(i)|
                within Cohort(i)

  value(i) is the SAME quantity the global percentile ranks — nothing is
  recomputed, only the field of competitors changes.

No new math.The cohort percentile uses the value computed globally — it is just ranked against fewer competitors. This keeps the pipeline fast, the storage cheap, and the output auditable. You can verify our cohort score by (a) looking up the item's value, (b) listing the cohort members, and (c) computing the rank yourself.

Category is the coarsest grouping ("wine," "bourbon," "single-origin coffee," etc.) and price is the list price at the time of most recent review. The ±20% band is symmetric: a $100 bottle's cohort is $80–$120.

Hotels are cohorted on location + price point rather than on a symmetric price band, because for a hotel the market is the city: a room in Manhattan and a room in Sultanahmet are not substitutes at any price. The cohort key is the city paired with a fixed nightly-rate band — $ under $90, $$ $90–169, $$$ $170–299, $$$$$300 and up — taken from the property's published rate range. Fixed thresholds, not per-city quartiles, so "$$" means the same thing everywhere and a band label doesn't silently redefine itself when the corpus grows. Hotels with no published rate form their own in-city band instead of being folded into a priced one.

Splitting a city by price makes some groups too small to rank against — a percentile over four peers is noise, not a measurement. So the hotel cohort widens until it has at least 12 members: city + price band first; failing that the whole city at every price point; failing that every hotel we rank in that price band worldwide. Each hotel page names the peer group it actually landed in and how many hotels were in it, so a widened cohort is never presented as a narrow one.

The AI-adjusted percentile — the third number on every detail page

Every detail page publishes three percentiles: the global one, the in-cohort one, and the AI-adjusted percentile, which answers a different question — what is left of this item's advantage once we discount it for how little evidence there is? This is the one figure on the site that is adjusted at all, and what it adjusts for is confidence: how much the sample behind a number actually supports it. It takes one form, identical in all six verticals: the score is shrunk toward the corpus average in proportion to how thin the sample is. Beside it, each item whose sample defines one also publishes the 90% confidence floor described above, which quantifies the same uncertainty from the interval side — a diagnostic that sets no percentile anywhere on the site.

Note what it does notdo: it never decides that one reviewer's opinion is worth more than another's. It is a statement about the sample, not about the people in it. And despite the name in the interface, it is not a model output — there is no training, no inference and no learned parameter anywhere in it, or anywhere else in this pipeline. It is one line of arithmetic, applied identically in all six verticals.

adj(i)         =  mean_z(i)  ·  n(i) / ( n(i) + K )        K = 53

score3(i)        =  100 · ECDF rank of adj(i) over the same published population
                    the global percentile ranks against
score3Cohort(i)  =  the same ECDF, taken inside item i's own cohort

  mean_z(i) = item i's mean reviewer z-score (the Step-2 output — the same
              quantity the global and in-cohort percentiles rank unshrunk)
  n(i)      = the number of calibrated reviewers that mean was taken over
  K         = 53, the median n across every ranked item on the site

The corpus average is zeroon the z-scale — that is what normalizing to each reviewer's own baseline means — so multiplying by n/(n+K) is literally pulling the item toward the average, hard when the sample is thin and barely at all when it is thick. An item with three glowing reviews keeps 3/(3+53) ≈ 5% of its measured distance from average; one with 3,000 keeps 98%. That is why a handful of raves cannot outrank thousands under this number.

Where K comes from. K = 53 is the median sample size across all 139,430 ranked items on the site, so by construction the shrink factor is exactly 0.5 at the median item: the typical listing is pulled precisely halfway to the average. Per-vertical medians differ a great deal — wines 108, cruises 42, hotels 34, movies 25, books 4, products 2 — and the pooled distribution runs 6 / 17 / 53 / 163 / 498 at the 10th, 25th, 50th, 75th and 90th percentiles. One constant is used everywhere rather than six, so the adjustment cannot be tuned per vertical after the fact.

Cross-check. Fitting the classical empirical-Bayes form — decomposing Var(mean_z | n) into a between-item component τ² and a sampling component σ²/n, then taking K = σ²/τ² — gives 2.5 for products, 4.7 for movies, 8.5 for wines, 27.1 for hotels and 32.3 for cruises. (It is undefined for books: at n = 3–6 the between-title variance is not separable from sampling noise, τ² ≤ 0.) K = 53 sits above the top of that range and well above its geometric mean of ~9.7, which means this constant shrinks harder than a fitted per-vertical prior would. That is the conservative direction for the claim the number makes.

Which n, per vertical.Wines, films & TV, books and cruises use the mean reviewer z from Step 2 and its qualifying-reviewer count. Hotels count only strictly calibrated reviewers. Products use the reviewer z over the wider admission rule described under Scope, and its review count. What differs between verticals is who qualifiesto be averaged, never how much any qualifying reviewer counts once they are in. In every case n is the reviewer count named directly under the chip, and on the verticals that publish a per-item statistics table the shrunk quantity itself is listed there as "sample-adjusted z".

Null, never zero.An item with nothing to shrink gets no figure: 462 hotels have no calibrated reviewer at all, and every product outside the z-ranked tier is excluded by exactly the same gate that excludes it from the global percentile. Those pages render an em dash. Publishing 0 would say "measured, and worst", which is a different and false claim. Across the site 139,430 items carry an adjusted percentile.

This figure does not replace the confidence floor published beside it, and the floor does not rank anything. The floor asks how low this item's quality could plausibly be given the sample; the adjusted percentile asks where the item ranks once its distance from average is discounted for sample size. They are two readings of the same underlying thinness, and both are published rather than blended — while the global percentile does neither, and reports what was measured. That division of labour is the whole point: you can see the measurement, see what the sample supports, and check the spread yourself, instead of being handed one number with the corrections already stirred in.

Scope — the six verticals this pipeline scores, and where coverage runs out

Everything above is the reviewer-z pipeline, and it is what produces the headline percentile for Films & TV, Hotels, Wines, Cruises, Books and — since 29 July 2026 — Products. Products previously shipped a different estimator entirely: a Bayesian-shrunk version of Amazon's own star average, under the score-schema id bayesian-site-v1. That is retired. Products now runs the same three steps as every other vertical, under reviewer-z-r3-mean-v2, with the wider reviewer-admission rule described below. A third id sits between the two: reviewer-z-r3-ci90-v1 shipped for part of 29 July 2026 and ranked the 90% CI floor of the same mean; it was superseded the same day by the mean itself, for the reasons set out under the products score. All three ids denote different scales and do not convert into one another — a percentile carrying an older id cannot be compared with one carrying the current one. What is still unique to Products is coverage: the reviewer overlap needed to normalize exists for only part of the catalog, and the part it does not reach is published without a score rather than with a substitute for one. Since 6 August 2026 that is literal — Amazon's star average and rating count are no longer published anywhere on the site, so an unscored product shows no rating at all.

Score schema by vertical. Every vertical is a percentile of the same statistic — an unweighted mean reviewer z-score. What differs is which reviewers qualify to enter that mean, and how much of the catalog it reaches. No vertical weights one reviewer above another.
Films & TV · Hotels · Wines · Cruises · BooksReviewer-z schema — the three steps above. Headline = empirical-CDF percentile of the unweighted mean reviewer z-score, in all five; hotels count only strictly calibrated reviewers into that mean, which is a rule about who qualifies, not about how the mean is ranked. Cohort = the same value re-ranked among peers (±20% price; city + price band for hotels).
Products (Amazon) — reviewer-z-r3-mean-v2The same schema with a wider admission rule: every reviewer with 2 or more ratings counts, including those whose ratings are all identical, whose dispersion is imputed from the source pool. They enter the mean at the same weight as everyone else. Headline = empirical-CDF percentile of the mean reviewer z. Cohort = the same mean re-ranked inside the product's category bucket, with no price band. 1,215,077 of 1,352,809 published products currently have enough reviewer overlap to carry a score; the other 137,732 are published with no rating figure at all — no percentile and no borrowed star average — and noindex.
Score schema by vertical. Every vertical is a percentile of the same statistic — an unweighted mean reviewer z-score. What differs is which reviewers qualify to enter that mean, and how much of the catalog it reaches. No vertical weights one reviewer above another.

What the products score actually is

The three steps above with the wider admission rule, then the same empirical-CDF step as everything else:

z(r)     =  ( rating(r) − μ_u )  /  σ_u          for each review r by reviewer u

  reviewers with n_u ≥ 2 count; n_u = 1 does not — a lone rating IS its own
  mean, so its z is identically 0 and carries no comparison
  σ_u = 0 (a reviewer whose 2+ ratings are identical) is admitted with an
  imputed σ̃ = 1.351, pooled from the dispersion of the σ_u > 0 reviewers.
  Such a reviewer rated the product AT their own average, so z(r) = 0 exactly:
  a genuine "this is par for me" vote. It moves mean_z toward 0 and adds to
  n(i) — it does not push the product up.

mean_z(i)   =  mean of z(r) over the reviews of product i     ← this is what ranks

score1(i)        =  100 · ECDF rank of mean_z(i) over the scored products
score1Cohort(i)  =  the same percentile recomputed within product i's category
score3(i)        =  100 · ECDF rank of mean_z(i) · n(i)/(n(i) + 53)

margin(i)   =  1.645 · sd(z over i) / √n(i)        90% one-tailed
ci_floor(i) =  mean_z(i) − margin(i)               shipped per product as
                                                   zMargin / zCiLower — still
                                                   computed, still published,
                                                   ranks nothing

The mean, not the floor, is the ranked quantity — and since this page is the specification, it is worth being exact about why that changed on 29 July 2026. Ranked from its bottom edge, a product measured by exactly one calibrated reviewer had no floor at all (a standard deviation over one value does not exist), and 884 more had a floor of zero width because their reviewers all landed on the same z. A rule meant to protect the thin end of the catalog was deleting it instead. It was also reordering on dispersion rather than on quality: on the shipped data a product at mean z +0.354 scored 59.9 while one at +0.500scored 55.0, purely because the second's reviewers disagreed more. So score1 now reports what was measured, the charge for sample size moved to the AI-adjusted percentile where it is a single legible multiplication, and the floor is still computed and still shipped on every product whose sample defines one — it simply orders nothing.

Amazon's star average and rating count are not published anywhere on this site.They were, as labelled reference figures beneath the score, until 6 August 2026. They are gone — not demoted. Rankquant collects them only as scrape diagnostics, because an impossible value in them is the fastest way to detect a parser reading the wrong part of a page. They never reach a product page, a data file, a sort order, a filter, or structured data. The only rating figure we publish is one we computed ourselves from individual reviewer-level ratings, and a product we have not measured carries no rating at all rather than borrowing the retailer's.

Why only part of the catalog is scored

A data limitation, stated plainly. Each ASIN's capture reads review page 1 only, and the crawl is sharded ASIN-disjoint — so a reviewer who rated two products in the catalog is often observed on only one of them. Of the 4,721,328 Amazon reviewers in the corpus, 1,594,196 rated two or more products and are therefore usable; the other 3,127,132 rated exactly one and are held back until they are seen again. That leaves 1,215,077 of 1,352,809 published products with at least one calibrated reviewer, and therefore a mean reviewer z to rank.

The other 137,732 keep their pages and their URLs, carry no rating figure of any kind, state that no normalized score exists yet, and are served noindex,follow and left out of the sitemap. They are not deleted and they are not back-filled with a substitute statistic: coverage rises every time the review capture goes deeper, and a product enters the scored set when its own evidence arrives.

Two rules that used to withhold scores here are gone with the CI-floor basis, and both were withholding real measurements. Products whose calibrated reviewers all landed on the same z are ranked now: a spread of zero is what those reviewers agreed on, and under a mean-ranked headline it is a measurement rather than a division by zero. Products measured by exactly one calibrated reviewer are ranked too — one z-scored review is thin, but it is evidence, and it is ours. What thinness still costs is indexing rather than a score: the 435,925 products standing on a single calibrated reviewer publish all three percentiles on their page and are still served noindex,follow and kept out of the sitemap, because one reviewer is one reviewer and that is not a page to ask Google to rank.

What we drop rather than rank

One rule, and it is about identifiers rather than ratings.

Until 6 August 2026 two further rules withheld roughly a quarter of the catalog, and it is worth saying plainly why they are gone, because the evidence behind them was sound and the conclusion was not. Both detected corruption in Amazon's own figures: an implausible 5.0 star average (the histogram spiked at 5.0 with 4,330 products while its neighbour at 4.9 held 648, and the reviews we actually scraped for those products mean 4.657), and rating counts repeated verbatim across unrelated products (87 carrying exactly 24,449; 38 carrying exactly 473,342 — a page-level total captured as one product's own).

Those were real defects and they mattered enormously when Amazon's figures werethe score. They stopped mattering when the score became reviewer-normalized and those figures stopped being published at all. A number nobody sees, that orders nothing and filters nothing, cannot mislead anyone by being wrong — and withholding a product because Amazon's number looks wrong hands Amazon's number the decision over what appears in our catalog, which is the same mistake in a less obvious costume. The detection still runs; it now tells us a scraper needs fixing instead of deciding what you get to see.

The rule the rest of the site follows is unchanged and is what governs here too: where the data cannot support the number, we do not publish the number. For a product with no calibrated reviewers that means no rating at all — not a percentile, and not Amazon's star average as a substitute.

Reading the spread — where taglines come from

The global, in-cohort and AI-adjusted percentiles are three views of the same number. When they agree, the item is simple to describe. When they diverge, the divergence itself is the story, and that is what our on-card taglines express. Every pattern below is a statement about peer group or sample size — the only two things that can move these three figures apart, precisely because no reviewer is weighted above another.

Spread patternWhat it meansTagline style
Global high, in-cohort much higherExceptional relative to its price peers; less dominant globally."Best in its price class; more moderate globally."
Global high, in-cohort lowerStrong overall but priced into a tough cohort."A strong global performer with fierce cohort competition."
Global high, AI-adjusted much lowerThe rating is high but thinly evidenced; the discount for sample size bites hard."Rated highly, but by very few people so far."
Global and AI-adjusted nearly equalDeep sample — there is almost nothing to discount."A well-established score, backed by a deep sample."
Global mid, in-cohort and AI-adjusted both highBeats its real competitors on solid evidence; the global field is simply broader."Quietly excellent against the things you would actually compare it to."
All three cluster tightly above 75Universally well-regarded, and well-evidenced."Exceptional by every lens we apply."

The decision tree above is the full tagline logic — there is no additional rule we apply off-page. No hidden editorial hand.

Where degrees of freedom enter

The Bessel-corrected nu − 1 in σu is degrees-of-freedom. A reviewer with nu = 2 has df = 1, and their σuis very noisy — but they still enter the z-score at full weight, because the alternative is deciding that some people's opinions count less. Nothing at the reviewer level corrects for it. The correction is at the item level, and it is the sample: reviewers with thin personal distributions produce noisier z, the only thing that damps that noise is more reviewers, and more reviewers is exactly what the AI-adjusted percentile pays for through n/(n+53). The published 90% floor registers the same thing from the other side, falling as the sample thins — and on products, where the margin is built from the observed spread of z rather than from N alone, falling further when the reviewers disagree. Thinness is charged for at the item level, never at the reviewer level.

The full treatment is at /theory/degrees-of-freedom/. The derivation of the published confidence floor and the choice of 90% (vs. 95% or 99%) is at /theory/confidence-intervals/.

Outbound links (fully disclosed)

There is no price-and-commission routing formula. We do not hold retailer price feeds, so nothing on this site compares prices across retailers or picks a destination by commission rate. Every outbound link is built deterministically from the record's own identifiers by one small module (lib/affiliate.ts), and each category gets a fixed destination:

No outbound link can influence the normalized score. The link is attached after the pipeline runs and is a pure function of identifiers the score never sees. The score is deterministic given the reviewer ratings and the published constants — there are no weights to buy, because there are no weights — and using the constants on this page you can reproduce or contest every number we publish.

Reproducibility

If you disagree with our output, you can check our work. Every step is published.

Editorial standard

The reference implementation is not currently distributed as a public repository, so this page is the specification: every constant the pipeline depends on — the nu ≥ 2 and σu> 0 admission rules, the shrinkage constant K = 53, the z = 1.645 behind the published 90% confidence floor, and the empirical-CDF percentile mapping — is written out above in full. There is no weight table among them, which is the point: the specification is complete because there is no editorial parameter left off it. Every detail page additionally shows its per-item intermediates (the mean reviewer z the headline ranks, the reviewer count, the sample-adjusted z behind the third figure, and — where the sample defines one — the 90% confidence floor), so you can audit or recompute any individual score from published numbers alone.

Frequently asked questions

What is the third percentile on every detail page?+
The AI-adjusted percentile — the same score, discounted for how much confidence the sample supports. Take the item's mean reviewer z-score, multiply it by n/(n+53) where n is its calibrated-reviewer count, and rank that shrunk value by empirical CDF against the same population the global percentile uses. That is the whole computation, and it is identical in all six verticals; the global and in-cohort percentiles rank the unshrunk mean itself, so this is the only figure of the three that adjusts anything. Because the corpus average is zero on the z-scale, multiplying by n/(n+53) pulls thin-sample items toward the average: an item with 3 reviewers keeps about 5% of its measured distance from average, one with 3,000 keeps 98%. K = 53 is the median sample size across all 139,430 ranked items, so the median item is pulled exactly halfway. The interface labels this chip "AI-adjusted percentile"; the computation is the single line of arithmetic above, with no model, no training and no inference in it. Items with nothing to shrink — 462 hotels with no calibrated reviewer, and every product outside the ranked tier — publish no figure rather than a zero.
Why normalize per reviewer rather than per source?+
Because reviewers, not publications, produce the scale. Two critics writing for the same publication can have wildly different personal ranges; a single critic reviewing both wine and bourbon uses one personal scale across both. Per-reviewer normalization is the finest grain that still has enough data — n_u ≥ 2 — to estimate a personal mean and standard deviation.
Is a product percentile the same number as a film percentile?+
It is now the same statistic, computed the same way: since 29 July 2026 products run the reviewer-z pipeline on this page (schema reviewer-z-r3-mean-v2) rather than the retired bayesian-site-v1. Both are the empirical-CDF percentile of an unweighted mean of per-reviewer normalized z-scores. They are still ranked against different populations — a product percentile is a rank among products — so we do not average them together, and any figure carrying an older schema id is on a different scale entirely. The real caveat for products is coverage: 1,215,077 of 1,352,809 listings have enough reviewer overlap to be scored, and the others carry no percentile at all.
Why the minimum σ_u > 0, not σ_u > some small number?+
Because any positive floor is arbitrary and noisy. A reviewer with σ_u = 0.2 is just as exploitable by the z-score formula as one with σ_u = 2. What we actually need to exclude is reviewers whose σ is mathematically undefined or zero (constant raters). In the verticals that admit them — Products — they enter with an imputed σ, and once admitted they count exactly as much as any other reviewer.
Why is the published confidence floor 90% and not 95%?+
A consumer review site is a different context from a drug trial — and note what this choice does and does not decide. The floor is a published diagnostic, not a ranking, so the confidence level sets how informative that figure is, never which item outranks which. At 95% the floor collapses toward zero for items with fewer than ~15 reviewers: every thin sample bottoms out together and the number stops telling a reader anything. At 90% the bound stays discriminating from about N ≥ 6 while still registering thin samples clearly. 90% one-tailed corresponds to z = 1.645; 95% would be z = 1.960. The choice is a defaults-matter call, not a mathematical one, and it's published and version-stable.
Why make cohort a re-ranking rather than a separate computation?+
Speed, storage, auditability. Re-ranking the same mean z within a cohort is O(k log k) per product; re-computing from scratch would require materializing a separate z-score table per cohort, which multiplies storage by the number of cohorts. The statistical information content is identical — a product's position within its cohort is fully determined by its own mean z and its cohort members' mean z.
What does ±20% price do at the edges?+
The symmetric ±20% band means a $100 product's cohort is $80–$120. For very cheap or very expensive products where the natural ±20% window produces thin cohorts, we fall back to the nearest-N rule: expand the band until the cohort contains at least 20 members. This edge-case rule is logged on the product page whenever invoked.
Can brands pay to rank higher?+
No, and there is not even a mechanism to buy. Outbound links are fixed per category and built from the record's own identifiers after the pipeline runs; they do not influence the mean reviewer z, any published confidence floor, or any percentile. Nor is there a source-weight or reviewer-weight table to lean on — every qualifying reviewer counts once and the same. The score is deterministic given the reviewer data and published constants.
Why include reviewers from other categories in the same reviewer pool?+
A person's rating scale is a property of the person, not the category. If a reviewer rates wine in a 6-point range and bourbon in a 4-point range, combining them gives a better estimate of their actual personal standard deviation than either alone. Category-specific pools would throw away information and shrink sample sizes for cross-category reviewers. Cross-category pooling assumes the person's rating calibration is somewhat stable across product types — an assumption we believe holds in the data but continue to validate.
How often are scores updated?+
When new reviews arrive the full pipeline re-runs. For products with active review volume that's a daily or weekly update. Historical scores remain visible so you can see how an evaluation has evolved. Every score on the site carries its computation date.
How is this different from IMDb's weighted average or Reddit's Wilson score?+
It's the same family, kept in a separate figure. IMDb's Top 250 shrinks toward a global prior; our AI-adjusted percentile does the same job with an explicit published constant, mean_z × n/(n+53) — and the retired products schema bayesian-site-v1 did that to Amazon's star average. Reddit's "best" sort uses a Wilson score lower bound on an up/down-vote binomial; we compute the analogous bound — a 90% one-tailed floor — and publish it alongside every score whose sample defines one, but nothing on the site is ranked by it. Two differences matter. First, the shrinkage is applied after per-reviewer normalization, so the quantity being protected is the right signal in the first place rather than a raw vote or rating. Second, it is published as its own percentile beside the unadjusted one instead of being blended into a single number a reader cannot take apart.
What if a product has no qualifying reviewers?+
We don't publish a normalized score. The product appears in our database as "not rated" with an explanation and, where possible, a pointer to the source reviews. We do not fabricate a score from thin data — that would defeat the entire point.