Research / Method
Audience simulations built on math, not LLMs
Consumer simulation today runs on language models and agent personas. Software asked to pretend to be your customer.
We have just proved there is a better way. We measured a real audience of 40,212 people, then fitted the mathematics that predicts what they do next: which trends spread, who adopts, and what they say when asked.
Published 27 August 2026. Every number below was scored against a future the model never saw.
How to read these. The last two are ranking skill: shown two people, how often does the model correctly say which one adopts first. 0.5 is a coin flip, 1.0 is perfect. Every score was measured on 7,151 held-out adoptions from the final stretch of the timeline, which the model was never shown.
fig. a — the audience graph, live · 18 real people from the 40,212-node german graph
every node is a real person · click to visit the account · hover for its numbersevery node is a real person · scroll the figure sideways · tap for its numbers, tap again to open the account
Same eighteen people, three ways of counting. By followers one account dwarfs the rest; under the measured lead score it shrinks, while an account with 99 followers leads on 596 adoption pairs. Reach is a claim, position is a theory, lead is a receipt.all 18 are real accounts; the portraits are illustrations · every edge is a confirmed follow · lead counted over the 207 person-to-person trends · position computed on all 60,015 edges
What you are looking at
Everything on this page comes from one real audience: 40,212 people in Germany, mostly students, measured on social media. We recorded who follows whom, and what each person did and when. We asked nobody anything.
Then we cut their timeline in two. The model learned from the earlier part and had to predict the later part, which it had never seen. That is what every number here is scored against: a stretch of real life that had already happened, with the answers held back.
One audience, shown end to end, including the parts that did not work. The figures below are drawn from that run rather than illustrated.
01 / The premise
You cannot simulate what you never measured
The standard approach is to prompt a language model into a persona and let a crowd of them stand in for your customers. Three things go wrong, and a better prompt fixes none of them.
- One brain. Ten thousand agents are one model wearing ten thousand names. When it is wrong, it is wrong the same way every time, so a thousand agents agreeing is not a thousand data points. It is one opinion, repeated.
- Personas flatten into stereotype. CoMPosT (Cheng, Piccardi and Yang, EMNLP 2023) measured two failures at once: the simulation loses the variation inside a group, and exaggerates whatever supposedly marks that group out. They call the result caricature, and found GPT-4 most susceptible to it when standing in for political and marginalised groups.
- Take away the ordering bias and the answers go random. Questioning the Survey Responses of Large Language Models (NeurIPS 2024) found model answers are governed by which option is labelled A. Randomise the order and models trend toward uniformly random responses, regardless of model size. Any apparent match to a demographic is then an artifact: a model resembles a group best when that group’s real answers happen to sit closest to uniform.
So we have a different approach. We start where the audience mapping ends: a real sample, connected by real follow edges, who actually knows whom. The German graph maps 40,212 people. We take its 830 most audience-central accounts and break their posts into pieces (person, did, thing, when, stance), 343,633 of them in one pass. No personas, no synthetic population. The people in our simulations exist.
02 / The ladder
Trends, then influence, then opinions — in that order
Three measurements, each standing on the one below: what is rising, who moves whom, what people will say.
What is actually rising — at three altitudes
Every piece is filed into a frozen 1,536-dimensional semantic map: 2,271 fine things → 567 themes → 75 domains. A candidate becomes a trend only by surviving trial — a sustained burst above its own baseline, beyond what noise could do, from multiple independent people. Newest run: 578 confirmed, 330 rejected.
Put plainly: we sort everything the audience talks about by meaning, from broad areas down to very specific things, then make each one prove it is genuinely rising before we call it a trend.
Who moves whom — measured, never assumed
The classical toolkit — PageRank, spectral centrality, structural entropy — runs as candidates only. The verdict is evidence: did your followers do the thing after you, beyond matched controls and a rewired graph? The payoff is a fitted probability of adoption per person × trend × week. That one equation is the simulation engine.
Put plainly: we do not take follower counts as influence. We check whether the people who follow you actually did the thing after you did it, and compare that against lookalikes who do not follow you.
What they will say when you ask
Each piece records how the person felt about the thing, so everyone in the graph accumulates a real opinion record. Ask the simulated audience a question and each answer is fitted from that person's own history first, then their crowd, then their network.
Put plainly: because we already know what each person has said about things like this before, we can answer a survey question the way they would, one person at a time.
fig. b — the trend atlas · the german audience’s 75 domains → 567 themes → 2,271 fine clusters, positioned by meaning
drag to rotate · switch altitude · click a cloud to isolate it · hover for namestap a cloud to isolate it
The real map, not an illustration. Distance is meaning, size is volume, blue embers are the court's 578 confirmations. The clouds are the domains and themes the dots actually form.cluster names are machine-generated; a few are noisy, and we left them that way rather than hand-edit the map
03 / The mathematics
Every trend must beat its own baseline
First, the trend test. Every topic has a normal daily level. To count as a trend, activity has to rise clearly above that level, stay there for days, and come from many different people — not from one loud account.
candidate a · 41 people · sustained burstconfirmed: trend
candidate b · 1 person · 31 postsrejected: one voice
Both spikes look like trends. Only one is. The left is 41 independent people sustaining a burst ~9× above the node's own baseline; the right is one account posting 31 times. Feed-counting tools report both as "trending".shapes drawn from real verdict patterns; day-level values illustrative
Second, the adoption model. For every person and every trend it estimates one thing: how likely is this person to pick this up this week? Move the sliders — the bars show how much each signal adds, and the curve turns the total into a probability.
the instrument, signal by signal — set the person, read the probability
the adoption instrument — fitted on this audience's own history
three measured signals in · one probability out · every prediction decomposes below
adopts this week
the whole model at once — every possible person on one sheet
This is the same model for every possible person at once. One direction: how many accounts they follow already adopted. The other: how hot the trend is. The higher and bluer the sheet, the more likely they adopt. The dot is the person you set with the sliders. Move your cursor over the sheet to try others, and click a segment to lock the reading in place.Tap a segment to lock the reading in place, or drag sideways across the sheet to try others. The exact weights are fitted per audience, and we keep them private.operating values illustrative; the fitted coefficients stay private
04 / What you can ask
Everything above exists so you can ask these four questions
Our model lets you question the audience directly, and every answer is built from what these people actually did, not from what a language model imagines they would say.
Field the survey before you field the survey
Answer distributions grounded in each person's real opinion record — not a persona's guess.
"How would 25–34s in this audience answer these ten questions?"Who would buy — and who would pass
Name the likeliest buyers, then compare them to the people who would pass — what the two groups do, follow and say differently.
"Who exactly are our 500 likeliest buyers — and how do they differ from the ones who won't?"What is rising, at which altitude, spreading where
Certified trends only, ranked by acceleration, with the accounts likeliest to carry each next.
"What is about to break out of the niche and go audience-wide?"Where to spend so it spreads
Point the seeding and creator budget at the people who provably move others — measured lead, not follower counts.
"Where do we put the launch budget so this actually spreads?"the finding that shapes all four answers
Behavior travels the graph.
What a person picks up next is written in who they follow — that is what takes adoption prediction from 58% to 81%, and it is why every answer above starts from the measured graph rather than from a persona.
05 / Same context, head to head
We gave the language model everything our model gets. It still came second.
We ran the fashionable alternative on the same German cases: a frontier LLM, handed the same context our own model works from. The same measured audience, the same network features, the same real adoption decisions, scored on the same held-out stretch. Nothing was withheld from it.
It came second, .78 against our .86. Then we ran it a third time with our measurements taken out, leaving it only a description of each person and what they seem to like. That is how consumer simulation is normally sold, and it scored .65, the worst of the three.
Same information in, three different answers out. Given everything our model is given, the language model still ranks the real adopters less accurately than the fitted mathematics does. Take the measured audience away and it drops to the bottom of the board, which is where persona-driven simulation actually operates.ranking skill on identical held-out cases from the German run · 0.5 is a coin flip
The reason is structural. An LLM predicts what a plausible person would do; a fitted model predicts what these specific people do, because it has watched them. Our predictions decompose into named, measured terms you can read.
Language models are not banished from this page. They keep two jobs: reading posts into those pieces, and grouping the pieces by meaning. What they never do is stand in for a person. After that step it is statistics: fitted, tested, falsifiable.
06 / Validation
We test every method on a world where we already know the answer
Early on, a model we trusted produced a confident wrong answer. Everything in this section exists so that cannot happen twice.
Before a method is allowed near a real audience, we build a fake one where we already know the truth, because we planted it. We decide in advance who influences whom and which trends spread, hand the method that invented world, and check whether it finds what we hid there. A method that cannot recover an answer we put in ourselves has no business estimating one we do not know.
We also write down what failure would look like before running the test rather than after, so we cannot talk ourselves into keeping something that did not work. More than fifteen of our own designs were killed that way.
Only then does a method meet real data, and there the test never changes. It learns from the first 80% of the audience’s timeline and is scored on the last 20%, which it has never seen. Every number on this page comes from that hidden stretch.
What this does not prove yet
- Prediction is not causation. People who follow each other also resemble each other. Our direct causal test stands at suggestive (p = .055) — published as exactly that.
- Retrospective so far. The standing forward forecast — freeze, wait, re-measure — is the next milestone, not a shipped one.
- One audience is one audience. A result on German students is a result on German students until the next audience replicates it.
by hugo
led by Artem Sergeyev
The offer
Bring an audience. We will show you who moves it.
A market, a segment, a competitor's customers. We build the graph, certify its trends, measure who moves whom, and hand you a simulation you can question — held-out scores attached. All of them, not just the flattering ones.
why we publish the misses
A simulation you cannot falsify is an opinion with a UI.
Everything above ships with its test: the planted-answer trials, the hidden 20%, the p-value that is not there yet. We publish the 330 rejections next to the 578 confirmations because the numbers do not need theatre.
Related reading: Mapping what every consumer audience actually wears · Consumer simulation, explained · Why AI consumer simulations fail for fashion brands · Hugo vs Aaru · Fashion demand forecasting