Methodology

What we do

For each product we collect the published expert reviews, lab measurements, manufacturer specifications and, where we can reach them, user-rating aggregates. We record what each source measured and concluded, exactly as published, and write a survey of where they agree and disagree. We do not test products ourselves and never present anyone else's testing as our own.

Which sources count

A product survey needs at least five distinct publishers, including at least two established expert outlets. Pages with fewer say "early consensus" and show lower confidence. We don't cite Wikipedia, content farms or other aggregators; they can lead us to a primary source, but the primary source is what we cite. Every page we cite is one we actually retrieved and read, and we record the date of the version we read, plus a link to an archived copy where one exists.

Publishers are graded by how rigorously they test: outlets that publish lab measurements carry the most weight, established review sites next, smaller or enthusiast outlets less, and retailer rating averages least. A manufacturer is never treated as a reviewer of its own product.

How the score is calculated

  1. Normalise. Reviewers publish on wildly different scales โ€” 4/5, 10/10, 87/100. Each rating is converted to a 0โ€“1 value. We never convert a rating a source didn't publish, and never invent one.
  2. Weight. Each rating is weighted by the publisher's tier, by how recent it is (weight halves roughly every eighteen months), and by the type of source. Retailer averages are additionally scaled by how many ratings they rest on.
  3. Shrink toward the category average. A product with two glowing reviews shouldn't outrank one with nine good ones, so every score is pulled toward the category's typical score by an amount that shrinks as real evidence accumulates.
  4. Report confidence. Fewer than three rated sources is "early consensus". Six or more, including at least two strong outlets, is a firm one.

The calculation is ordinary code with tests, not a judgement by a language model, so the same evidence always produces the same number. Every product page can show the full table of inputs behind its score.

What the score leaves out

Some outlets publish a written verdict but lock their scores and measurements behind an account. We cite their words and leave their numbers out rather than guess at them, which means a rigorous source can inform a page's text while contributing nothing to its score. Reviews without any published rating work the same way.

How pages are written

Research, drafting and checking are done by AI systems following the rules in oureditorial policy, with a separate checking pass that re-opens every cited page and tests each claim against it. A person reviews and approves every page before it is published. When we get something wrong, we fix it and record it on thecorrections page.