Methodology
| Hotel | Dan | Ari | Mel | Jen | Tim | Avg |
|---|---|---|---|---|---|---|
| A | 5/5 | — | 5/5 | 3/5 | 3/5 | 4.0 |
| B | 5/5 | 5/5 | 5/5 | 3/5 | 3/5 | 4.2 |
| C | — | — | — | 4/5 | 4/5 | 4.0 |
| D | 4/5 | 5/5 | — | 1/5 | 2/5 | 3.0 |
| Irrelevant | Most important information | |||||
Summary
raw score average
Follow Hotel B: ringed at 4.2, first on the raw average — and underlined at 49, third once every reviewer is re-centred.
Other sites would have you choose Hotel B. Rankquant clearly shows you math-perfected Hotel C is superior.
Who is doing this, and why the math
Rankquant is a team of statistically-trained editors. We are not a review site that hired a data scientist; we are applied statisticians who believe the way review scores are published today is straightforwardly broken. Our job on this page is to explain, in plain English paired with the exact equations, how we turn messy real-world reviews into a number you can actually trust.
None of this math is novel. Per-reviewer normalization, reviewer fixed-effects aggregation, and shrinking a thin sample toward the population mean are textbook techniques in psychometrics, meta-analysis, and sports analytics — see /theory/founding-metrics for a first-principles tour of the five primitives this pipeline is built from. What's different is publishing them, applied rigorously, as the basis of a consumer review site.
The problem every review site has
Every major review surface suffers severe right-skew inflation. Raw averages cannot distinguish a genuinely exceptional product from a merely average one.
Amazon's long-run average across billions of reviews across all categories.
Marketplace Pulse review analyses
Yelp's long-run category-averaged restaurant rating.
Yelp transparency reports
Booking.com's long-run hotel rating; hotels below 8.0 are flagged as low-rated.
Booking.com scoring bands
When everything is "excellent," the word carries no information. And there's a second problem hiding underneath the inflation: different reviewers use the same numbers to mean different things. One reviewer's 88 is another reviewer's ceiling; one's 3-stars is another's "I liked it." Averaging raw numbers across reviewers mixes those private scales together and throws away signal.
Our fix has two moves. First, normalize every reviewer onto their own z-scale so personal grading habits wash out. Second, count every one of those reviewers equally and let the size of the sample — not our opinion of the reviewer — decide how much confidence the result earns.
The method, in three steps
Step 1 — per-reviewer z-score normalization
For each reviewer u, compute the reviewer's personal mean μu and standard deviation σu over all of their reviews across all products in our database. Their rating of product i is then converted to a z-score:
z_{u,i} = ( r_{u,i} − μ_u ) / σ_u
where
r_{u,i} = reviewer u's raw rating of product i
μ_u = reviewer u's personal mean rating across all their reviews
σ_u = reviewer u's personal standard deviation (Bessel-corrected, df = n_u − 1)
n_u = number of reviews reviewer u has written in our datasetTwo reviewers who disagree on scale but agree on quality produce identical z-scores. A reviewer who always rates 90–96 and a reviewer who always rates 75–88 will both produce z ≈ +1.3for their personal favorites — because that's where each reviewer sits relative to their own distribution. The z-score is dimensionless. It's the unit reviewers actually share.
Who counts as a "qualifying reviewer"?
Not every reviewer's opinions belong in the z-scale. Two conditions:
- nu ≥ 2. One review per reviewer contains no personal distribution — μ_u and σ_u are undefined.
- σu > 0. A reviewer who has given every product the same rating carries no signal about relative quality. Including them amounts to dividing by zero.
A reviewer with nu ≥ 2 but σu = 0isn't useless — they just can't be normalized the standard way. We keep them on file, and in the verticals that admit them (Products does — see Scope) we impute a plausible σ from their source's pooled dispersion. When they are admitted they enter the same unweighted mean as everyone else, at exactly the same weight.
Cross-category pooling. A reviewer who writes both wine and bourbon reviews goes into one reviewer pool with one μu and one σucomputed across everything they've rated. That's deliberate: a person's personal scale is a property of the person, not the product category. Pooling keeps nu large and σu stable.
Step 2 — combine the reviewers, weighting none of them
After Step 1, every item has a list of z-scores — one per qualifying reviewer. We take their plain arithmetic mean. That is the whole of Step 2.
mean_z(i) = (1 / N_i) · Σ_{u ∈ Q_i} z_{u,i} ← equal weight per reviewer
Q_i = qualifying reviewers of item i (n_u ≥ 2, σ_u > 0)
N_i = |Q_i|Every qualifying reviewer counts the same.An amateur with five reviews counts exactly as much as a professional with five thousand. There is no credibility weight, no source weight, no per-reviewer multiplier and no editorial thumb anywhere on the scale. The only thing that changes a reviewer's influence is how many other reviewers are in the average with them.
This is a deliberate constraint rather than a missing feature. The moment a weight exists, the interesting question stops being "what did the reviewers say" and becomes "who decided the weights, and what would change if they were different" — and that question is unanswerable from outside. A single unweighted mean is the version of this number a reader can reconstruct from published data without trusting our judgment about anybody.
We do not sell weight, and we do not assign it either. Every qualifying reviewer counts once.
Earlier versions of this page described a second, source-weighted aggregate and a third "broadened" one, presented as three lenses. Those were designed but never shipped: no published score has ever depended on a source weight, and there has never been a live weight table. The description has been removed rather than left standing as a claim about a thing that does not exist. What Step 3 does publish is three rankings of this one number — not three different numbers.
Step 3 — rank that one number, three ways
Step 2 produced a single quantity per item. Step 3 does not produce more quantities — it ranks that one quantity three times, and those three rankings are the three percentiles on every detail page.
| Global percentile | The empirical-CDF rank of the mean z-score against every item we rank in that vertical, expressed 0–100. This is the headline. It is built from the mean z-score and nothing else — no confidence floor, no shrinkage. It reports the measurement as made. |
|---|---|
| In-cohort percentile | The same mean z-score, re-ranked inside a narrower peer group — same category and a close price band, or city and nightly-rate band for hotels. Same quantity, smaller field. A re-ranking, not a second computation. |
| AI-adjusted percentile | The same mean z-score after it is discounted for how much confidence the sample supports: multiplied by n/(n+53), so thin samples are pulled toward the corpus average. One line of arithmetic, applied identically in all six verticals — no model, no training, no inference. This is the only one of the three that adjusts anything. |
Charging a score for its uncertainty
This is the part that matters most for what you see on a page — and it happens in exactly one of the three figures. An item with 4 reviewers averaging z = +2.1 might be exceptional, or its mean might be noise; one with 80 reviewers averaging z = +1.6 has paid its statistical dues. The global and in-cohort percentiles do not arbitrate that — they rank the mean as measured, and +2.1 goes ahead of +1.6. The arbitration lives in the AI-adjusted percentile, which discounts each item in proportion to how little evidence stands behind it. Note where the penalty lands: on the sample, never on the reviewer. We never decide that a particular person's opinion is worth less; we decide how much confidence a given number of opinions buys.
The discount is one multiplication — mean_z · n/(n+53) — applied identically in all six verticals, and it is written out in full under the AI-adjusted percentile below. Separately, and ordering nothing, we publish the lower bound of a 90% one-tailed confidence intervalaround each item's mean, wherever the sample defines one, as a diagnostic. That floor answers a question no percentile does: "given this sample size, what is a defensibly pessimistic estimate of the item's true quality?"
Ẑ = mean_z(i), the unweighted mean from Step 2
SE(Ẑ) = 1 / √N (standard error of the mean)
floor = Ẑ − 1.645 · SE(Ẑ) ← one-tailed 90% CI lower bound
N = the number of qualifying reviewers in the mean. Because every reviewer
carries the same weight, the effective sample size IS N — there is no
weighted correction to apply.
Published per item — ciFloorZ in the catalog exports, zCiLower and zMargin
on products — and read next to the score. Undefined at N = 1, where it is
absent rather than faked. No percentile on the site is ranked by it.Worked example — two real-world shapes, and what each figure does with them:
| Product | N (reviewers) | Mean Ẑ — what global ranks | n/(n+53) | Adjusted z — what AI-adjusted ranks | 90% floor (published, ranks nothing) |
|---|---|---|---|---|---|
| Thin-sample darling | 4 | +2.10 | 0.070 | +0.15 | +1.28 |
| Well-reviewed consensus | 80 | +1.60 | 0.602 | +0.96 | +1.42 |
On the mean — which is what the global and in-cohort percentiles rank — the darling wins, +2.10 to +1.60, and we publish that rather than quietly correcting it. On the AI-adjusted percentile the order reverses, +0.96 to +0.15, because four reviewers keep only 7% of their measured distance from average while eighty keep 60%. The published confidence floors, +1.28 and +1.42, point the same way without ordering anything. This is by design, and it's the same intuition Bayesian sports ratings, IMDb's Top 250 formula, and Wilson score confidence intervals (used by Reddit and Yelp internally) all use: penalize uncertainty, reward consistency — kept in its own figure instead of folded silently into the headline.
From a z-score to a 0–100 percentile
A mean z-score is fine for math but means nothing to a reader. So we rank every item's value against every other item's and express the result as a percentile using the empirical CDF:
p_global(i) = 100 · rank( value(i) ) / N_total
value(i) = mean_z(i) in all six verticals — the measurement itself,
with no confidence floor and no shrinkage in it
rank() = ordinal rank, ties split at midpoint
N_total = number of items with a valid value in that vertical
An item with no mean_z is absent from the denominator and stays unranked. It
is never parked at 0 or at the corpus average to fill the column.A percentile of 90 means the item scores higher than 90% of everything else we rank in that vertical. 50 is the median. 10 means it sits in the bottom 10%.
Cohort percentiles — a re-ranking, not a re-computation
Global percentile answers "how does this product compare to everything we measure?" — which is useful but sometimes unfair. A $14 bottle that beats all other $14 bottles is doing exactly what a $14 bottle should do; showing it at the 25th global percentile (against $200 Burgundy) hides that achievement.
So we also publish a cohort percentile. It is the same value, re-ranked within a narrower peer group.
Cohort(i) = { j : category(j) = category(i)
AND |price(j) − price(i)| / price(i) ≤ 0.20 }
p_cohort(i) = 100 · rank( value(i) ) / |Cohort(i)|
within Cohort(i)
value(i) is the SAME quantity the global percentile ranks — nothing is
recomputed, only the field of competitors changes.No new math.The cohort percentile uses the value computed globally — it is just ranked against fewer competitors. This keeps the pipeline fast, the storage cheap, and the output auditable. You can verify our cohort score by (a) looking up the item's value, (b) listing the cohort members, and (c) computing the rank yourself.
Category is the coarsest grouping ("wine," "bourbon," "single-origin coffee," etc.) and price is the list price at the time of most recent review. The ±20% band is symmetric: a $100 bottle's cohort is $80–$120.
Hotels are cohorted on location + price point rather than on a symmetric price band, because for a hotel the market is the city: a room in Manhattan and a room in Sultanahmet are not substitutes at any price. The cohort key is the city paired with a fixed nightly-rate band — $ under $90, $$ $90–169, $$$ $170–299, $$$$$300 and up — taken from the property's published rate range. Fixed thresholds, not per-city quartiles, so "$$" means the same thing everywhere and a band label doesn't silently redefine itself when the corpus grows. Hotels with no published rate form their own in-city band instead of being folded into a priced one.
Splitting a city by price makes some groups too small to rank against — a percentile over four peers is noise, not a measurement. So the hotel cohort widens until it has at least 12 members: city + price band first; failing that the whole city at every price point; failing that every hotel we rank in that price band worldwide. Each hotel page names the peer group it actually landed in and how many hotels were in it, so a widened cohort is never presented as a narrow one.
The AI-adjusted percentile — the third number on every detail page
Every detail page publishes three percentiles: the global one, the in-cohort one, and the AI-adjusted percentile, which answers a different question — what is left of this item's advantage once we discount it for how little evidence there is? This is the one figure on the site that is adjusted at all, and what it adjusts for is confidence: how much the sample behind a number actually supports it. It takes one form, identical in all six verticals: the score is shrunk toward the corpus average in proportion to how thin the sample is. Beside it, each item whose sample defines one also publishes the 90% confidence floor described above, which quantifies the same uncertainty from the interval side — a diagnostic that sets no percentile anywhere on the site.
Note what it does notdo: it never decides that one reviewer's opinion is worth more than another's. It is a statement about the sample, not about the people in it. And despite the name in the interface, it is not a model output — there is no training, no inference and no learned parameter anywhere in it, or anywhere else in this pipeline. It is one line of arithmetic, applied identically in all six verticals.
adj(i) = mean_z(i) · n(i) / ( n(i) + K ) K = 53
score3(i) = 100 · ECDF rank of adj(i) over the same published population
the global percentile ranks against
score3Cohort(i) = the same ECDF, taken inside item i's own cohort
mean_z(i) = item i's mean reviewer z-score (the Step-2 output — the same
quantity the global and in-cohort percentiles rank unshrunk)
n(i) = the number of calibrated reviewers that mean was taken over
K = 53, the median n across every ranked item on the siteThe corpus average is zeroon the z-scale — that is what normalizing to each reviewer's own baseline means — so multiplying by n/(n+K) is literally pulling the item toward the average, hard when the sample is thin and barely at all when it is thick. An item with three glowing reviews keeps 3/(3+53) ≈ 5% of its measured distance from average; one with 3,000 keeps 98%. That is why a handful of raves cannot outrank thousands under this number.
Where K comes from. K = 53 is the median sample size across all 139,430 ranked items on the site, so by construction the shrink factor is exactly 0.5 at the median item: the typical listing is pulled precisely halfway to the average. Per-vertical medians differ a great deal — wines 108, cruises 42, hotels 34, movies 25, books 4, products 2 — and the pooled distribution runs 6 / 17 / 53 / 163 / 498 at the 10th, 25th, 50th, 75th and 90th percentiles. One constant is used everywhere rather than six, so the adjustment cannot be tuned per vertical after the fact.
Cross-check. Fitting the classical empirical-Bayes form — decomposing Var(mean_z | n) into a between-item component τ² and a sampling component σ²/n, then taking K = σ²/τ² — gives 2.5 for products, 4.7 for movies, 8.5 for wines, 27.1 for hotels and 32.3 for cruises. (It is undefined for books: at n = 3–6 the between-title variance is not separable from sampling noise, τ² ≤ 0.) K = 53 sits above the top of that range and well above its geometric mean of ~9.7, which means this constant shrinks harder than a fitted per-vertical prior would. That is the conservative direction for the claim the number makes.
Which n, per vertical.Wines, films & TV, books and cruises use the mean reviewer z from Step 2 and its qualifying-reviewer count. Hotels count only strictly calibrated reviewers. Products use the reviewer z over the wider admission rule described under Scope, and its review count. What differs between verticals is who qualifiesto be averaged, never how much any qualifying reviewer counts once they are in. In every case n is the reviewer count named directly under the chip, and on the verticals that publish a per-item statistics table the shrunk quantity itself is listed there as "sample-adjusted z".
Null, never zero.An item with nothing to shrink gets no figure: 462 hotels have no calibrated reviewer at all, and every product outside the z-ranked tier is excluded by exactly the same gate that excludes it from the global percentile. Those pages render an em dash. Publishing 0 would say "measured, and worst", which is a different and false claim. Across the site 139,430 items carry an adjusted percentile.
This figure does not replace the confidence floor published beside it, and the floor does not rank anything. The floor asks how low this item's quality could plausibly be given the sample; the adjusted percentile asks where the item ranks once its distance from average is discounted for sample size. They are two readings of the same underlying thinness, and both are published rather than blended — while the global percentile does neither, and reports what was measured. That division of labour is the whole point: you can see the measurement, see what the sample supports, and check the spread yourself, instead of being handed one number with the corrections already stirred in.
Scope — the six verticals this pipeline scores, and where coverage runs out
Everything above is the reviewer-z pipeline, and it is what produces the headline percentile for Films & TV, Hotels, Wines, Cruises, Books and — since 29 July 2026 — Products. Products previously shipped a different estimator entirely: a Bayesian-shrunk version of Amazon's own star average, under the score-schema id bayesian-site-v1. That is retired. Products now runs the same three steps as every other vertical, under reviewer-z-r3-mean-v2, with the wider reviewer-admission rule described below. A third id sits between the two: reviewer-z-r3-ci90-v1 shipped for part of 29 July 2026 and ranked the 90% CI floor of the same mean; it was superseded the same day by the mean itself, for the reasons set out under the products score. All three ids denote different scales and do not convert into one another — a percentile carrying an older id cannot be compared with one carrying the current one. What is still unique to Products is coverage: the reviewer overlap needed to normalize exists for only part of the catalog, and the part it does not reach is published without a score rather than with a substitute for one. Since 6 August 2026 that is literal — Amazon's star average and rating count are no longer published anywhere on the site, so an unscored product shows no rating at all.
| Films & TV · Hotels · Wines · Cruises · Books | Reviewer-z schema — the three steps above. Headline = empirical-CDF percentile of the unweighted mean reviewer z-score, in all five; hotels count only strictly calibrated reviewers into that mean, which is a rule about who qualifies, not about how the mean is ranked. Cohort = the same value re-ranked among peers (±20% price; city + price band for hotels). |
|---|---|
| Products (Amazon) — reviewer-z-r3-mean-v2 | The same schema with a wider admission rule: every reviewer with 2 or more ratings counts, including those whose ratings are all identical, whose dispersion is imputed from the source pool. They enter the mean at the same weight as everyone else. Headline = empirical-CDF percentile of the mean reviewer z. Cohort = the same mean re-ranked inside the product's category bucket, with no price band. 1,215,077 of 1,352,809 published products currently have enough reviewer overlap to carry a score; the other 137,732 are published with no rating figure at all — no percentile and no borrowed star average — and noindex. |
What the products score actually is
The three steps above with the wider admission rule, then the same empirical-CDF step as everything else:
z(r) = ( rating(r) − μ_u ) / σ_u for each review r by reviewer u
reviewers with n_u ≥ 2 count; n_u = 1 does not — a lone rating IS its own
mean, so its z is identically 0 and carries no comparison
σ_u = 0 (a reviewer whose 2+ ratings are identical) is admitted with an
imputed σ̃ = 1.351, pooled from the dispersion of the σ_u > 0 reviewers.
Such a reviewer rated the product AT their own average, so z(r) = 0 exactly:
a genuine "this is par for me" vote. It moves mean_z toward 0 and adds to
n(i) — it does not push the product up.
mean_z(i) = mean of z(r) over the reviews of product i ← this is what ranks
score1(i) = 100 · ECDF rank of mean_z(i) over the scored products
score1Cohort(i) = the same percentile recomputed within product i's category
score3(i) = 100 · ECDF rank of mean_z(i) · n(i)/(n(i) + 53)
margin(i) = 1.645 · sd(z over i) / √n(i) 90% one-tailed
ci_floor(i) = mean_z(i) − margin(i) shipped per product as
zMargin / zCiLower — still
computed, still published,
ranks nothingThe mean, not the floor, is the ranked quantity — and since this page is the specification, it is worth being exact about why that changed on 29 July 2026. Ranked from its bottom edge, a product measured by exactly one calibrated reviewer had no floor at all (a standard deviation over one value does not exist), and 884 more had a floor of zero width because their reviewers all landed on the same z. A rule meant to protect the thin end of the catalog was deleting it instead. It was also reordering on dispersion rather than on quality: on the shipped data a product at mean z +0.354 scored 59.9 while one at +0.500scored 55.0, purely because the second's reviewers disagreed more. So score1 now reports what was measured, the charge for sample size moved to the AI-adjusted percentile where it is a single legible multiplication, and the floor is still computed and still shipped on every product whose sample defines one — it simply orders nothing.
Amazon's star average and rating count are not published anywhere on this site.They were, as labelled reference figures beneath the score, until 6 August 2026. They are gone — not demoted. Rankquant collects them only as scrape diagnostics, because an impossible value in them is the fastest way to detect a parser reading the wrong part of a page. They never reach a product page, a data file, a sort order, a filter, or structured data. The only rating figure we publish is one we computed ourselves from individual reviewer-level ratings, and a product we have not measured carries no rating at all rather than borrowing the retailer's.
Why only part of the catalog is scored
A data limitation, stated plainly. Each ASIN's capture reads review page 1 only, and the crawl is sharded ASIN-disjoint — so a reviewer who rated two products in the catalog is often observed on only one of them. Of the 4,721,328 Amazon reviewers in the corpus, 1,594,196 rated two or more products and are therefore usable; the other 3,127,132 rated exactly one and are held back until they are seen again. That leaves 1,215,077 of 1,352,809 published products with at least one calibrated reviewer, and therefore a mean reviewer z to rank.
The other 137,732 keep their pages and their URLs, carry no rating figure of any kind, state that no normalized score exists yet, and are served noindex,follow and left out of the sitemap. They are not deleted and they are not back-filled with a substitute statistic: coverage rises every time the review capture goes deeper, and a product enters the scored set when its own evidence arrives.
Two rules that used to withhold scores here are gone with the CI-floor basis, and both were withholding real measurements. Products whose calibrated reviewers all landed on the same z are ranked now: a spread of zero is what those reviewers agreed on, and under a mean-ranked headline it is a measurement rather than a division by zero. Products measured by exactly one calibrated reviewer are ranked too — one z-scored review is thin, but it is evidence, and it is ours. What thinness still costs is indexing rather than a score: the 435,925 products standing on a single calibrated reviewer publish all three percentiles on their page and are still served noindex,follow and kept out of the sitemap, because one reviewer is one reviewer and that is not a page to ask Google to rank.
What we drop rather than rank
One rule, and it is about identifiers rather than ratings.
- ISBN-namespace ASINs (a 10-character identifier that does not begin with
B). Those are ISBNs, not Amazon product ASINs: they identify published media, which Rankquant already ranks in its Books catalog. Amazon's own metadata cannot separate them — all of them arrived labelled Office Products, Toys or Electronics and not one as Books — so the identifier namespace is the test. Left in, the Portuguese edition of The Tipping Point, filed under Electronics, sorted to the top of the entire products catalog.
Until 6 August 2026 two further rules withheld roughly a quarter of the catalog, and it is worth saying plainly why they are gone, because the evidence behind them was sound and the conclusion was not. Both detected corruption in Amazon's own figures: an implausible 5.0 star average (the histogram spiked at 5.0 with 4,330 products while its neighbour at 4.9 held 648, and the reviews we actually scraped for those products mean 4.657), and rating counts repeated verbatim across unrelated products (87 carrying exactly 24,449; 38 carrying exactly 473,342 — a page-level total captured as one product's own).
Those were real defects and they mattered enormously when Amazon's figures werethe score. They stopped mattering when the score became reviewer-normalized and those figures stopped being published at all. A number nobody sees, that orders nothing and filters nothing, cannot mislead anyone by being wrong — and withholding a product because Amazon's number looks wrong hands Amazon's number the decision over what appears in our catalog, which is the same mistake in a less obvious costume. The detection still runs; it now tells us a scraper needs fixing instead of deciding what you get to see.
The rule the rest of the site follows is unchanged and is what governs here too: where the data cannot support the number, we do not publish the number. For a product with no calibrated reviewers that means no rating at all — not a percentile, and not Amazon's star average as a substitute.
Reading the spread — where taglines come from
The global, in-cohort and AI-adjusted percentiles are three views of the same number. When they agree, the item is simple to describe. When they diverge, the divergence itself is the story, and that is what our on-card taglines express. Every pattern below is a statement about peer group or sample size — the only two things that can move these three figures apart, precisely because no reviewer is weighted above another.
| Spread pattern | What it means | Tagline style |
|---|---|---|
| Global high, in-cohort much higher | Exceptional relative to its price peers; less dominant globally. | "Best in its price class; more moderate globally." |
| Global high, in-cohort lower | Strong overall but priced into a tough cohort. | "A strong global performer with fierce cohort competition." |
| Global high, AI-adjusted much lower | The rating is high but thinly evidenced; the discount for sample size bites hard. | "Rated highly, but by very few people so far." |
| Global and AI-adjusted nearly equal | Deep sample — there is almost nothing to discount. | "A well-established score, backed by a deep sample." |
| Global mid, in-cohort and AI-adjusted both high | Beats its real competitors on solid evidence; the global field is simply broader. | "Quietly excellent against the things you would actually compare it to." |
| All three cluster tightly above 75 | Universally well-regarded, and well-evidenced. | "Exceptional by every lens we apply." |
The decision tree above is the full tagline logic — there is no additional rule we apply off-page. No hidden editorial hand.
Where degrees of freedom enter
The Bessel-corrected nu − 1 in σu is degrees-of-freedom. A reviewer with nu = 2 has df = 1, and their σuis very noisy — but they still enter the z-score at full weight, because the alternative is deciding that some people's opinions count less. Nothing at the reviewer level corrects for it. The correction is at the item level, and it is the sample: reviewers with thin personal distributions produce noisier z, the only thing that damps that noise is more reviewers, and more reviewers is exactly what the AI-adjusted percentile pays for through n/(n+53). The published 90% floor registers the same thing from the other side, falling as the sample thins — and on products, where the margin is built from the observed spread of z rather than from N alone, falling further when the reviewers disagree. Thinness is charged for at the item level, never at the reviewer level.
The full treatment is at /theory/degrees-of-freedom/. The derivation of the published confidence floor and the choice of 90% (vs. 95% or 99%) is at /theory/confidence-intervals/.
Outbound links (fully disclosed)
There is no price-and-commission routing formula. We do not hold retailer price feeds, so nothing on this site compares prices across retailers or picks a destination by commission rate. Every outbound link is built deterministically from the record's own identifiers by one small module (lib/affiliate.ts), and each category gets a fixed destination:
- Wines— the wine's Vivino page when we hold a Vivino ID, plus a Wine-Searcher query built from winery + name + vintage. Both untracked; they earn nothing. There is no wine affiliate program to join, but Amazon.com sells a small slice of the catalog, and those bottles additionally get an Amazon Associates link (see the footer disclosure). A wine qualifies only on an exactwinery + name match against our Amazon title corpus — a “Reserve” or single-vineyard bottling is a different wine and is never substituted for the one you searched. The match is computed from identifiers alone, before the score exists.
- Products— the Amazon detail page for the record's ASIN. This is the largest linked surface on the site, one link per published product. An Amazon Associates tracked link (see the footer disclosure); it may earn a commission at no additional cost to you. The link is built from the ASIN alone, so it is fixed before the score exists and cannot vary by commission rate.
- Books— the Amazon detail page for the record's ISBN, when it has one. Formatted as an Amazon Associates link (see the footer disclosure).
- Films & TV — a Prime Video search for title + year, shown in the detail panel on /movies/catalog/. Also formatted as an Associates link. Individual film pages carry no streaming link.
- Cruises — a CruiseDirectlink, tracked through Commission Junction, on every ship page. Where CruiseDirect publishes a page for that specific ship the link goes straight to it; for the lines they do not carry it falls back to a search on the ship's name. This one does earn: 3% of the fare, and only once the passenger has actually completed the sailing — a booking on its own pays nothing. The Cruise Critic and Booking.com links beside it are untracked and earn nothing.
- Hotels — a Booking.com search built from the property name. Untracked; we earn nothing.
No outbound link can influence the normalized score. The link is attached after the pipeline runs and is a pure function of identifiers the score never sees. The score is deterministic given the reviewer ratings and the published constants — there are no weights to buy, because there are no weights — and using the constants on this page you can reproduce or contest every number we publish.
Reproducibility
If you disagree with our output, you can check our work. Every step is published.
The reference implementation is not currently distributed as a public repository, so this page is the specification: every constant the pipeline depends on — the nu ≥ 2 and σu> 0 admission rules, the shrinkage constant K = 53, the z = 1.645 behind the published 90% confidence floor, and the empirical-CDF percentile mapping — is written out above in full. There is no weight table among them, which is the point: the specification is complete because there is no editorial parameter left off it. Every detail page additionally shows its per-item intermediates (the mean reviewer z the headline ranks, the reviewer count, the sample-adjusted z behind the third figure, and — where the sample defines one — the 90% confidence floor), so you can audit or recompute any individual score from published numbers alone.