Method · Engine version 8

How the score is computed

SureVett grades how trustworthy a product's reviews look. It does not grade how good the product is. A superb product can carry a manipulated review section, and an ordinary product can have completely honest reviews and score well here. Everything below is the actual method, at the thresholds the shipping extension uses.

Most review checkers hand you a number and ask you to trust it. That is the thing we set out not to do, so this page is the whole calculation: the five signals, what each one rewards and punishes, the exact cutoffs, when we refuse to grade at all, and the cases the method is known to miss. If you disagree with a grade after reading this, you will be able to say precisely which part you disagree with. That is the point.

The composite

Five signals are computed from the public data on the product page. Each returns a score from 0 to 100, and the grade is their weighted average. The analysis runs in your browser; nothing about the reviews is sent anywhere.

The first number is the weight as declared in the engine; the second is the share it actually carries when all five signals are present. They differ because the declared weights sum to 0.80 rather than 1.00. A sixth signal held the missing 0.20 until we retired it in August 2026: Amazon stopped publishing the figure it depended on, so it fired zero times in more than 1,500 analyses while holding the second-highest weight in the score. We removed it rather than keep advertising a measurement that could not happen.

Missing signals are excluded, never scored as a neutral 50. When a product page does not publish what a signal needs, the remaining signals are reweighted to sum to 1 and the panel tells you how many were available. A fake neutral would quietly drag every grade toward the middle and hide the gap; this way the reasoning stays visible and confidence drops instead.

The grade bands

You can see how these come out across everything we have ever graded on the live grade distribution on the homepage. It is the real one, refreshed daily, not a flattering sample. A checker that returns roughly the same score for every product could not publish its distribution, which is why we publish ours.

Signal 1: rating distribution

The heaviest signal, and the J-curve test. Authentic Amazon review sections are lopsided in a particular way: a lot of five stars, a real minority of one stars, and a thin but populated middle. Faked ones tend to be either a near-uniform wall of five stars or a distribution with no negatives at all.

The base score is a smooth curve over five-star concentration rather than a set of buckets, so distributions that differ get scores that differ. It peaks around 65 to 72 percent five-star, which is where genuinely well-liked products sit, and falls off in both directions: at 85 percent it is 58, at 90 percent it is 40, at 95 percent it is 22, and a 100 percent five-star wall scores 5. Below about 45 percent it eases down too, since that is a poorly-reviewed product rather than a manipulated one.

Three adjustments then apply.

A bonus for a real J-shape. Between 45 and 96 percent five-star, a one-star share between 2 and 20 percent adds 2 points, and a populated middle (at least 2 percent three-star and 3 percent four-star) adds 1 more.

A cap for the classic fake signature. Zero one-star reviews alongside 85 percent or more five-star caps the score at 15; an entirely empty middle at that concentration caps it at 20. These only apply below 2,000 reviews, because Amazon rounds the percentages it displays. On a 24,000-review product, a displayed zero can be hiding a hundred real one-star reviews, so a zero is only evidence when the sample is small enough to make it a true zero.

Two floors that pull back from accusing. Above 80 percent five-star, a product that still carries a real critical tail (at least 2 percent one-star and 3 percent across two-to-four) earns a floor that scales with how substantial that tail is, from 45 up to 78. Concentration on its own is not evidence of manipulation; concentration without a negative tail is. Separately, a large sample at 95 percent five-star or below floors at 40 above 5,000 reviews and 45 above 20,000, because faking a distribution gets more expensive roughly in proportion to its size. A flat wall above 95 percent stays harsh at any sample size.

Signal 2: review velocity

Reviews divided by the age of the listing. Organic reviews trickle in; purchased ones arrive in batches, because that is how they are delivered.

Three launch-spike patterns score hardest: more than 200 reviews in under 30 days scores 10, more than 500 in under 60 days scores 15, and more than 10 reviews a day inside the first 90 days scores 25. Otherwise the score follows the daily rate: under 0.3 a day scores 90, under 1 scores 80, under 3 scores 70, under 8 scores 55, under 15 scores 35, and anything faster scores 20.

This is the signal most often unavailable, by a wide margin. It needs a listing date, and Amazon publishes a first-available date on only about half of US product pages and far fewer outside the US. Across our analyses to date, roughly 80 percent report product age as unknown, at which point velocity is dropped and the remaining four carry the grade. When you see that row marked unavailable, it is Amazon's omission rather than a failure on our side, and we would rather say so than guess an age.

Signal 3: review count

Sample size, scored on its own because a small sample makes every other signal less reliable. Under 5 reviews scores 25, under 15 scores 35, under 50 scores 50, under 200 scores 65, under 1,000 scores 75, under 5,000 scores 85, and above that 90.

Signal 4: seller trust

Who is actually selling the item. Sold by Amazon scores 95; sold by a third party but fulfilled by Amazon scores 75; a third-party seller not using FBA scores 40; and an unresolvable seller scores 50. This is the signal a bug once broke badly, misreading the seller on Amazon's newer offer layout and corrupting roughly half of all analyses until we found and fixed it in August 2026.

Signal 5: brand and product signals

Four sub-checks, scored as points earned over points available, so a page that publishes only some of them is judged on what it publishes rather than penalised for the rest.

Brand name, up to 3 points. A name that reads as gibberish, meaning a run of four or more letters with no vowel at all, earns nothing, and so does a name over 30 characters that also contains digits. A merely long name earns 1. Everything else earns the full 3. All-caps names are not penalised, because YETI, SONY, ASUS, LEGO and IKEA exist. The vowel check only runs on ASCII names, so that BJÖRN or a Japanese brand on amazon.co.jp is not stripped to a vowel-less husk and falsely flagged.

Best Sellers Rank, up to 3 points. Top 1,000 earns 3, top 10,000 earns 2, top 100,000 earns 1, below that nothing. Having a rank at all means Amazon is tracking real sales.

Answered questions, up to 2 points. 50 or more earns 2, 10 or more earns 1. Customer questions are hard to fake at volume and indicate a real audience.

Listing age, up to 2 points. Over a year earns 2, over 90 days earns 1.

When we refuse to grade

A letter is not always the honest output, so three states exist below it.

Fewer than 10 reviews and no grade is shown at all, because there is not enough there to say anything. Between 10 and 24 reviews the grade appears with a limited-data caveat attached. Zero reviews returns nothing to analyse. And if the page cannot be read at all, which happens when Amazon changes a layout, you get an explicit read-failure state with a way to report it, rather than a confident grade computed from nothing.

Confidence is reported alongside the grade and is a direct function of how many signals had data: five is high, three or four is medium, fewer is low.

What this method does not catch

Any tool claiming to catch every fake review in 2026 is selling something. AI-written reviews are good enough now that linguistic tells are unreliable, which is a large part of why the shipping engine is statistical and deterministic rather than an LLM.

The scoring is deliberately biased toward missing fakes rather than accusing real products. Those floors on the rating-distribution signal exist because accusing an honest product is the error a shopper sees and is harmed by, while missing a manipulated listing that has learned to seed a few one-star reviews is the error nobody notices. That trade is deliberate, and it has a cost: a seller sophisticated enough to plant a realistic critical tail can buy cover under those same floors. We keep a test fixture pinned to exactly that blind spot so it stays visible to us.

The method also cannot see anything Amazon does not publish. It does not read review text, which means it cannot judge whether an individual review is fake, only whether the pattern across a listing looks shaped. It cannot tell a manipulated listing from a genuinely viral one on velocity alone. And it cannot connect the same factory product sold under a dozen different brand names, which is where the fraud in low-cost categories has largely moved.

We surface signals, not certainty. A grade is a reason to look closer, not a verdict.

Version and changes

This describes engine version 8. The version is recorded with every grade, and cached results computed by an older version are recalculated rather than served, so a scoring change reaches every product rather than only new ones. The numbers on this page are transcribed from the shipping engine and are covered by a golden-master test, so if they ever drift from the code, the tests fail before the page is wrong.

Tired of fake reviews? Try SureVett, it's free.

Add to Chrome — Free