There is a version of this page that just says "highly accurate" next to a logo wall. The problem with the whole AI research category is that the output is fluent enough to be convincing whether or not it is right, so an unquantified accuracy claim is worth nothing. The only useful answer is a number, a method, and a list of the things it gets wrong.
71% winner accuracy · 0.039 mean JSD · 70 held-out questions · 19 audiences · 11 countries · 5 languages
What was actually tested
Hugo builds an audience from real consumer signal and can then simulate how that audience would answer a question. To test whether the simulation is any good, we ran it against held-out real survey data: questions with known, published human answers that the system had not seen.
Two metrics, because they measure different things.
- Winner accuracy (71%): how often the simulated audience picked the same top answer as the real one. This is the number that maps to a decision, because most commercial questions come down to "which of these do they prefer".
- Mean Jensen-Shannon divergence (0.039): how closely the whole distribution of answers matched, not just the ranking. JSD runs 0 (identical) to 1 (completely different), so 0.039 is a tight match. This is the stricter test: you can get the winner right by luck and still have the shape badly wrong.
Reporting only the first number would flatter us. Reporting only the second would hide whether the answer is actionable. Both matter.
Where it is strongest
Accuracy is not uniform across question types, and pretending otherwise is how people get burned. Four areas came out clearly ahead:
- Brand and product taste. The core commercial use case, and the best-performing one: on the US beauty audiences, 100% of favourite-brand winners were correct across fragrance, skincare, makeup and retailer.
- Behaviour legible from content. Gaming habits, platform usage, exercise. Things people demonstrate publicly rather than self-report.
- Attitudes, values and beliefs. Vote intention, societal outlook, religion, mental-health prevalence.
- Generalising across geography and language. A Finnish-seeded Finland audience scored 0.028 JSD, better than the overall mean and on par with US audiences.
That last one is worth dwelling on, because it is the usual failure point for anything built on social signal. Methods that depend on the density of English-language conversation fall apart outside the big markets. This one did not.
Where it breaks
Three failure modes showed up consistently. We publish them because a benchmark without a failure list is marketing.
1. Off-content options
An audience built from a single seed channel over-indexes that channel when you ask about cross-platform behaviour. Ask a TikTok-seeded audience how much time they spend on TikTok and you will get an answer biased by construction.
2. Niche posted behaviour
Communities formed around one interest over-represent that interest when reporting day-to-day habits. A running community will tell you people run more than they do, because running is the reason the community exists.
3. Flattened magnitudes
The most important one for decisions. The model tends to pick the right number one while understating how dominant it is, and it defaults attitude scales toward the middle. So a real 60/40 split may come back as 52/48.
The practical rule that falls out: trust the ranking more than the gap. Use it to decide which option wins. Do not use it to size the margin, and do not put the simulated percentage in a business case as if it were an incidence figure.
How this differs from synthetic users
Worth separating, because the two get conflated and one of them benchmarks badly. A recent systematic review of synthetic participants across 182 studies found accuracy close to a coin flip.
The distinction is generated versus retrieved. A synthetic user is asked to invent a plausible response, and a model producing a plausible response converges on the most conventional answer for the persona described. It is therefore right when the answer is obvious and wrong exactly when the answer is surprising, which is the worst possible error distribution for research.
Hugo's audiences are built from retrieved signal: real statements real people made, which is why the simulation can be checked against them and why the benchmark above is meaningful at all. It is also why the everyday product output carries verbatims and citations rather than just a percentage. The related discipline on the retrieval side is covered in the run where 78% of what we gathered was discarded.
How to pressure-test any accuracy claim
Whoever you are evaluating, including us, four questions separate a real benchmark from a brochure:
- Was the test data held out? If the system saw the answers, the score means nothing.
- What is the distribution metric, not just the hit rate? Winner accuracy alone hides the shape.
- How many questions, audiences and markets? A benchmark on one country and ten questions is an anecdote.
- What is on the failure list? A vendor who cannot name where their system breaks has not looked.
The honest summary
On 70 held-out questions across 19 audiences in 11 countries, Hugo picked the right top answer 71% of the time with distributions matching to 0.039 JSD. It is sharpest on brand and product taste, which is also where most commercial decisions live. It understates margins, over-indexes seed channels, and should not be used to size a market.
This is v1, run on gpt-4.1-mini in June 2026. The full per-audience and per-question breakdown, including the individual questions it got wrong, is on our validation page.