Blog / Method

How accurate is AI consumer research? Here are our numbers

"Is it accurate?" is the only question that matters in this category, and almost nobody selling into it publishes a number. Here is ours, measured against real survey data on held-out questions, including the three places it breaks.

By the Hugo team · Updated August 6, 2026 · 8 min read

There is a version of this page that just says "highly accurate" next to a logo wall. The problem with the whole AI research category is that the output is fluent enough to be convincing whether or not it is right, so an unquantified accuracy claim is worth nothing. The only useful answer is a number, a method, and a list of the things it gets wrong.

71% winner accuracy · 0.039 mean JSD · 70 held-out questions · 19 audiences · 11 countries · 5 languages

What was actually tested

Hugo builds an audience from real consumer signal and can then simulate how that audience would answer a question. To test whether the simulation is any good, we ran it against held-out real survey data: questions with known, published human answers that the system had not seen.

Two metrics, because they measure different things.

Reporting only the first number would flatter us. Reporting only the second would hide whether the answer is actionable. Both matter.

Where it is strongest

Accuracy is not uniform across question types, and pretending otherwise is how people get burned. Four areas came out clearly ahead:

That last one is worth dwelling on, because it is the usual failure point for anything built on social signal. Methods that depend on the density of English-language conversation fall apart outside the big markets. This one did not.

Where it breaks

Three failure modes showed up consistently. We publish them because a benchmark without a failure list is marketing.

1. Off-content options

An audience built from a single seed channel over-indexes that channel when you ask about cross-platform behaviour. Ask a TikTok-seeded audience how much time they spend on TikTok and you will get an answer biased by construction.

2. Niche posted behaviour

Communities formed around one interest over-represent that interest when reporting day-to-day habits. A running community will tell you people run more than they do, because running is the reason the community exists.

3. Flattened magnitudes

The most important one for decisions. The model tends to pick the right number one while understating how dominant it is, and it defaults attitude scales toward the middle. So a real 60/40 split may come back as 52/48.

The practical rule that falls out: trust the ranking more than the gap. Use it to decide which option wins. Do not use it to size the margin, and do not put the simulated percentage in a business case as if it were an incidence figure.

How this differs from synthetic users

Worth separating, because the two get conflated and one of them benchmarks badly. A recent systematic review of synthetic participants across 182 studies found accuracy close to a coin flip.

The distinction is generated versus retrieved. A synthetic user is asked to invent a plausible response, and a model producing a plausible response converges on the most conventional answer for the persona described. It is therefore right when the answer is obvious and wrong exactly when the answer is surprising, which is the worst possible error distribution for research.

Hugo's audiences are built from retrieved signal: real statements real people made, which is why the simulation can be checked against them and why the benchmark above is meaningful at all. It is also why the everyday product output carries verbatims and citations rather than just a percentage. The related discipline on the retrieval side is covered in the run where 78% of what we gathered was discarded.

How to pressure-test any accuracy claim

Whoever you are evaluating, including us, four questions separate a real benchmark from a brochure:

The honest summary

On 70 held-out questions across 19 audiences in 11 countries, Hugo picked the right top answer 71% of the time with distributions matching to 0.039 JSD. It is sharpest on brand and product taste, which is also where most commercial decisions live. It understates margins, over-indexes seed channels, and should not be used to size a market.

This is v1, run on gpt-4.1-mini in June 2026. The full per-audience and per-question breakdown, including the individual questions it got wrong, is on our validation page.

Test it against something you already know the answer to.

The fastest way to judge this is to ask about an audience you understand deeply and check where the answer disagrees with you. Bring one to the call.