Review and rating data
Ratings, review counts and star distributions across products and sellers, collected on a schedule so you can watch them move. The metrics that tell you what customers think, without republishing what customers wrote.
Metrics rather than copied text
Review bodies are written by individuals and published under a platform's terms. Collecting them wholesale creates copyright and personal-data questions that most buyers do not want attached to a dataset, and in practice the raw text is rarely what drives a decision anyway.
What does drive decisions is the shape of the feedback. A product holding 4.6 across three thousand reviews is a different proposition from one holding 4.6 across nine. A rating sliding from 4.4 to 3.9 over two months is a quality problem in progress. A star distribution with a heavy one-star tail behaves nothing like a flat one, even at the same average.
So we collect the structured signal by default. If your use case genuinely requires review text — a licensed research project, or a platform where you hold rights — we will discuss it, with the legal position established before any collection starts.
What each row contains
| Field | Example | Notes |
|---|---|---|
| sku | BH-1729841 | Product identifier on the source site |
| title | 35mm f/1.8 prime lens | |
| brand | Sony | Enables brand-level sentiment comparison |
| rating_avg | 4.62 | As displayed by the retailer |
| rating_count | 1,284 | The number an average is worth judging against |
| rating_dist | {"5":812,"4":301,"3":92,"2":41,"1":38} | Star breakdown where the site exposes it |
| verified_share | 0.87 | Proportion marked as verified purchases, where available |
| answered_questions | 34 | Q&A volume, a useful proxy for pre-purchase friction |
| seller_rating | 4.8 | Marketplace seller score, separate from the product |
| seller_rating_count | 15,402 | |
| collected_at | 2026-09-08T06:00:12Z | Every reading timestamped, so trends are reconstructable |
What the numbers are good for
Catching a quality problem early
A rating that starts sliding usually moves weeks before returns data reflects it. Tracked weekly across your own SKUs, this is one of the cheapest early warning signals available.
Benchmarking against rivals
Your 4.3 means one thing in a category where everyone sits at 4.5 and something else entirely where the average is 3.8. Category-wide collection gives you the baseline.
Ranking a category by traction
Review count is the closest public proxy for sales volume on most retail sites. Sorted by count within a category, it shows which products actually move.
Vetting marketplace sellers
Seller scores and volumes across the marketplaces you sell on, useful for competitive positioning and for screening resellers of your own brand.
Two things worth knowing about review data
Ratings are aggregated differently on every site
Some retailers show a straight mean, some weight recent reviews more heavily, some syndicate reviews across product variants so a colour with nine reviews displays the parent's three thousand. Comparing raw averages between sites without accounting for this produces conclusions that are confidently wrong. We document the aggregation behaviour of each source in the delivery note.
The trend needs a starting point
A single collection gives you a snapshot: where things stand today. The genuinely valuable output — direction of travel — only exists once you have several runs behind you. If you are considering this, starting the feed early costs little and is the only way to have history when you need it.
What it costs
Quoted per source and interval. Review metrics are usually collected alongside price or catalog data on the same run, which makes adding them to an existing feed considerably cheaper than commissioning them separately.
What we do not collect
No review text by default, no reviewer names, no profile links, no photographs submitted by customers. The output is numeric and aggregate. This keeps the dataset clear of personal data and of the copyright that attaches to individual reviews.
Review data, in practice
Why not just collect the review text?
Individual reviews are authored works published under platform terms, and they often contain personal detail. Redistributing them creates copyright and privacy exposure for whoever holds the file. The structured metrics answer most commercial questions without any of that, so they are our default.
Can you do sentiment analysis on the reviews?
Not as a standard product, since it requires the text. Where a client has rights to the content — their own reviews exported from their own platform, for example — we can process it. For competitor sentiment, the rating distribution over time carries most of the same signal.
How often should review data be collected?
Weekly suits most catalogs. Review counts accumulate slowly, so daily collection multiplies cost without adding much resolution. The exception is a launch window, where daily readings for the first few weeks are genuinely informative.
Can you detect fake reviews?
We do not sell that as a finding, because it is a judgement rather than a fact. What we deliver are the inputs people use to make that judgement: sudden jumps in review count, unusual distribution shapes, a low verified-purchase share. What you conclude from those is yours.
Does this work on marketplaces?
Yes, with the caveat that the largest marketplaces are also the most heavily defended, which affects cost and reliability. We scope those case by case and are straightforward about which ones we can run at a sensible price.
Other services
Start tracking before you need the history
Send a category and the products that matter. We will run a free sample so you can see the metrics on your own market.