Every benchmark run, newest first. Each run has its own permalink.
This page validates Hugo Simulations, the poll feature that asks a simulated panel a question and returns a distribution of answers. It does not cover Hugo's research feature, which is evaluated separately.
Every release of the simulator is benchmarked against published real-world surveys it has never seen. We update this page with every run.
Headline metric
Across 70 held-out questions covering 19 audiences in 11 countries, Hugo gets the top answer right 71% of the time and the full distribution matches reality to within 0.039 JSD on average. Hugo tells you what an audience will pick and how strongly they lean, and it's sharpest where it matters most for commercial decisions: brand and product taste, and behavior you can read off content.
A simple yes/no per question: did Hugo pick the same #1 answer as the real survey? Only looks at the top choice.
Shown on the headline above, in the Scorecard (questions Hugo got right per audience), and on each question in the All questions section as a green check or red x.
Jensen–Shannon divergence: how different the full answer distributions are. 0 = identical, 1 = opposite. Under 0.1 is a tight match.
Got right only checks the winner; JSD compares the whole shape, so Hugo can miss the #1 but still land close if the split is similar, or get #1 right but be off on the rest.
A second view: how far each audience lands from reality on the 0 → 1 distribution scale
As close as two real surveys would differ
Directionally tight
Right shape, soft on magnitudes
Trends only
Don't trust
Audiences from the 19-audience run, best to worst
| Audience | Account responses | Questions | JSD | Got right |
|---|---|---|---|---|
Brazilian evangelical youth Brazil · faith-driven Gen-Z | 109 | 2 | 0.010 | 2/2 |
US Gen-Z gamers United States · ages 13–25 | 61 | 4 | 0.012 | 3/4 |
Swedish climate youth Sweden · climate-engaged | 105 | 4 | 0.021 | 2/4 |
US outdoor & hiking enthusiasts United States · outdoors community | 102 | 3 | 0.024 | 1/3 |
Brazilian fitness women Brazil · body-positive fitness | 95 | 3 | 0.026 | 3/3 |
Swedish Gen-Z Sweden · 16–25 | 70 | 12 | 0.026 | 8/12 |
US Gen-Z retail investors United States · personal finance · 18–29 | 63 | 3 | 0.028 | 2/3 |
Finnish Gen-Z Finland · 16–30 | 60 | 4 | 0.028 | 3/4 |
| Mean | 70 | 0.039 | 50/70 |
What it gets right, where it's still developing
Real vs simulated, per audience
For each audience, Hugo builds a panel of 50–500 simulated personas from public social content and has them answer the survey question. We then compare Hugo's answers to a real, published survey.
Each persona is grounded in a public account's actual activity, what they post, what they engage with, the language they use. Hugo receives that profile and answers the way the person plausibly would.
Hugo never sees the real numbers. Every survey is verified against its original source and linked on each question, so the comparison can be audited end-to-end.
Audiences are chosen to stress-test the system across geographies, languages, and interests, from US gamers and beauty shoppers to Kenyan Gen Z, Brazilian fitness women, and Finnish students.
Most of these surveys were published after the model's training cutoff, so Hugo can't have memorized them. Recall doesn't explain the accuracy.
Tested on 70 held-out questions across 19 audiences in 11 countries. Hugo tells you what an audience will pick and how strongly they lean, in minutes instead of weeks.
Run your own
Bring a survey you trust. We'll run the same held-out test on your audience, live, in the demo.
Book a demo