Compare / Aaru
Hugo vs Aaru
Aaru's own site says people are poor predictors of their own behaviour. We agree with that, and it is the only thing the two companies agree on. Aaru assembles a synthetic population out of census columns, anonymised card transactions, point of interest visits and search demand, then runs a scenario through it. We find the real people, learn a model of each one from what that person actually posted and wore, and simulate how influence travels between them. No census column holds a product nobody has manufactured yet. A person writing "why does nobody make this in wool" holds it exactly.
Published 24 August 2026. Our account of Aaru's method comes from their own public materials rather than from a demo, so read it as a careful reading and not an audit.
Short answer
Choose Aaru if your question is macro and the behaviour is already written down somewhere: how a well documented population splits on a choice that already exists, what a policy does, how a message lands across a broad demographic spread. Reconstructing a population from records is a reasonable way to get at that, and the EY wealth study recreation is a real number against a real benchmark rather than a testimonial.
Choose Hugo for anything a consumer brand actually decides. What to make next, how a launch will spread, which campaign travels and which one dies inside the audience that already follows you, what a price change does to the segments you cannot afford to lose. We answer those from a map of real people and a model of the influence between them, not from a population assembled out of census columns.
The difference in one line: a synthetic population can only contain what somebody already recorded. A real one contains what people want next, because they are saying it right now.
01 / Common ground
The disagreement is not about whether to use AI
Both companies start from the same observation, and it is correct. People are unreliable narrators of their own behaviour. A survey measures what someone is willing to say about themselves to a stranger who is paying them, which is a different quantity from what they will do in a shop at 9pm.
So the argument is not AI against traditional research. Both of us think the panel, the focus group and the six week agency engagement are the wrong instrument for most commercial questions. The argument is about what you build the replacement out of. Aaru builds it out of records of what populations did. We build it out of observation of what real people are doing and saying right now, and out of models we trained ourselves on top of that. Everything below follows from that one choice.
02 / Why AI simulations fail
A language model cannot play a consumer
Start with the version sold under the name synthetic users: a large language model handed a paragraph that says 34 year old nurse, Malmö, two children, 34,000 kronor a month, and then asked which of three jackets she buys. Run it ten thousand times with ten thousand paragraphs and you have a synthetic panel. It does not work, and the five reasons are structural rather than fixable with a better prompt.
Trained to be rational. Consumers are not.
A language model is trained, and then tuned again, to produce the defensible answer. It weighs options like an expected-utility calculator and explains itself with reasons. Real buying runs on habit, mood, status, envy, timing and who was seen wearing it first, with the reason invented afterwards. Ask a model why someone bought the jacket and it will construct a rational account of a decision that was never rational, because that is the one thing it is best at.
Ten thousand agents share one brain
Running a large population of agents on one foundation model does not give you a population. It gives you one brain wearing ten thousand name tags: the same weights, the same latent space, the same post-training. Their errors are correlated, not independent, so averaging them does not cancel a shared bias. It launders it into a confident number with a false sense of sample size.
Alignment deletes the mechanism
Models are deliberately tuned to suppress status judgement, in-group signalling, gendered assumption and class inference. Those are not defects when you are modelling a shopper. They are most of the mechanism. A model built to refuse to be prejudiced is a poor instrument for measuring prejudice, and buying clothes is full of it.
The corpus is the wrong sample
The corpus is Wikipedia, news archives, published books, documentation and argument: text written by people who stopped to think. The closest thing to unfiltered consumer instinct is a Reddit thread, a product review or a TikTok comment, and that was always a thin slice of the training data. Since 2023 the platforms have licensed or locked most of it, so the slice is getting thinner, not thicker.
You cannot honestly backtest it
The obvious check is to replay a historical study and see whether the simulation reproduces it. Almost every published study you would use is already in the training data, so a good score measures recall rather than prediction. Anyone claiming a clean historical validation of an LLM population should be asked exactly how they ruled leakage out.
Right when it is obvious, wrong when it matters
All five push the same way. The model returns the most conventional answer for the persona described, so it is correct when the answer was already guessable and wrong exactly when the finding would have been worth paying for. That is the worst possible error distribution for research, and it is why we did not build on it.
What we built instead03 / The evidence
This is not our opinion. It is the published record.
Simulated humans have been studied hard for three years now, and the failure modes have names. Nobody has to take our word for the argument above. If you want the plain explainer first, what a consumer simulation is and how one works covers the ground before this page starts arguing.
- The largest review found they mimic rather than predict. A systematic literature review of synthetic participants across 182 studies, spanning psychology, economics, marketing, healthcare and HCI, sorted the problems into four classes: cognitive misalignment, distortion, misleading believability and contamination. Its conclusion casts strong doubt on the ability of synthetic users to do more than mimic data that was already collected.
- They answer surveys nothing like people do. A NeurIPS 2024 paper, Questioning the Survey Responses of Large Language Models, put census questions to models and found the answers dominated by ordering bias, with entropy that stays high and flat no matter what is asked. Real populations get more certain on some questions and less on others. Models do not.
- Distinct personas collapse into one. This is the one brain problem, measured. Give a thousand agents a thousand different profiles and they converge on a narrow band of behaviour anyway, and the 2026 work on persona collapse finds that pushing fidelity up makes it worse rather than better, because the fidelity is bought with stereotype: gender or social class ends up dominating the variance.
- Caricature, not a person. The CoMPosT framework measured this directly and named it. Simulated personas come out as flattened caricatures that miss how many directions a real person points in at once.
- The variance is fake, and so is the confidence. Models reproduce mean responses reasonably and then show artificially low variance and systematic overconfidence around them. One study found silicon-sampled responses inflating affective polarisation by roughly seven times against the human benchmark.
- Pollsters make the sharpest version of the point. An AI poll collects no new data. It is a model of what a poll would have shown, not a measurement of what anyone thinks, and when the results are cut by demographic subgroup the gaps against real polling have run to 10 to 20 points. That is fine for a model and disqualifying for a measurement, and the two keep getting reported as the same thing.
A simulation trained on what was already collected cannot tell you the thing nobody has collected yet. That is not a bug in the current generation. It is the definition of the method.
04 / What Aaru actually does
Their method is better than the strawman, so argue with the real one
It would be convenient for us if Aaru were nothing but prompted personas. They are not, and a comparison page that pretends otherwise is worthless. Aaru was founded in 2024, has raised at a reported billion dollar headline valuation, and counts Accenture as an investor and partner. The method they describe publicly is closer to statistics than to role play, and the statistical half of it is good.
They start from public and licensed records of how populations live, work and spend: census and labour data, anonymised transaction records, point of interest visits, search demand, media use, plus whatever a client brings in segments and purchase histories. Sources are weighted against each other, so agreement strengthens a signal and disagreement flags it.
The part worth respecting is the joint modelling. A population is not a spreadsheet of independent columns. Age and work status are not independent, and neither are sex and industry. Rather than treating each source row as an agent, they learn the network of relationships the data supports and sample complete profiles out of it, checking that the relationships hold not only pairwise but in three way and higher combinations. Training runs at two resolutions at once: the population totals, and how an individual response varies by person and condition, with a check that the individual responses add back up to the observed totals. EY reported a median Spearman correlation of 0.90 recreating its 2025 global wealth study across 3,600 affluent investors.
If the answer to your question is already latent in recorded behaviour, reconstructing the joint distribution is a sound way to get at it.
One thing their public materials do not describe in the same detail is the response step. The population construction is documented. How an agent then answers is not, while press coverage of their polling work describes thousands of AI respondents returning a result in under two minutes. So the strongest available reading is that records build the population and a model produces the responses, which puts everything in section 02 back on the table for the second half of the pipeline no matter how good the first half is.
The disagreement that does not depend on any of that starts one step earlier, at what the records contain.
05 / The structural limit
A synthetic population can only contain what somebody already recorded
This is not a criticism of the engineering. It is a property of the input, and it applies to every simulation built this way.
- Every source is a record of a choice among things that already exist. A transaction is a purchase of something on sale. A search query names something already nameable. A census column was defined years before the question you are asking.
- Demand for a product nobody makes has no column. There is no transaction for the jacket that was never manufactured. That demand exists as a sentence, in a comment, and it reads like "why does nobody make this in a tall fit" or "I would buy this instantly if it came in wool". This is the class of finding no record set can contain, and it is the one that changes a range plan.
- Taste is not a demographic property. Two people with identical rows in every register want opposite things, because they watch different people. Income and postcode do not carry that. What carries it is who someone is connected to, and connection is not in the records at all.
- Records lag, and a trend is what happens before they catch up. By the time a shift is visible in transaction data it is visible to every competitor buying the same transaction data. A joint distribution rebuilt from records is a high resolution picture of the recent past.
- The irrational part is the product, not the noise. In fashion most of the decision is status, timing and who wore it first. Averaged into a population distribution that disappears, and the distribution is exactly where a record-based simulation is strongest and a brand's decision is least served.
None of this makes record-based simulation useless. It makes it a rear facing instrument with very high resolution. That is genuinely valuable when the question is macro. It is the wrong shape entirely when the question is what should exist next, which is the question a consumer brand asks every week.
06 / What we built instead
Observe the real population, then model it properly
We do not start with a model of a person and hope it resembles one. We start with real people, and then train our own models on what those people actually did. No step in that chain asks a language model to be a shopper.
Reach real people, not the ones an algorithm surfaces
Search a hashtag like stockholm style and you get the accounts a ranking algorithm decided perform Stockholm best, not the people who live there, and the selection probability is set by a system nobody outside the platform can see. We reach people by chain referral through their real connections instead, the method survey science calls respondent driven sampling. How that works.
See what they actually wear and buy
Computer vision reads garments out of ordinary video rather than out of outfit posts. In our New York map, 2% of frames were posts about clothing. The other 98% is someone filming a kitchen, a dog or a night out while wearing clothes, which is the closest thing to an unposed sample of what people actually own. 832 distinct garments came off 876 accounts that way. The full run.
Hear what they ask for and cannot find
Requests, complaints, comparisons and the moment somebody gives up, in their own words, on your products and your competitors'. "Why does nobody make this in a tall fit." "I would buy this instantly if it came in wool." "Bought the cheaper one, the zip went in a month." That is demand for something that does not exist yet, a product flaw that never reached your support inbox, and a switch you were never told about, and no record set anywhere contains any of the three.
Learn a model per person, not a persona
Each person in the map carries a small model trained on what that person actually posted, wore and reacted to. It is a vector, not a paragraph of English, and it outputs a predicted reaction directly. Because it works in that space, disliking one thing carries over to the near neighbours of that thing without anyone writing a rule for it. No step in the chain reasons about what a person like this would probably say.
Connect them on a graph
A node is a person or a trend. An edge is a person engaging with a trend, or two trends appearing together. That graph is the object we model, because it holds the one thing no record set contains: who is downstream of whom, and how far a thing travels after the first hundred people see it.
Throw most of it away
In one published run we gathered 1,328 candidate posts and kept 287. The discarded 78% were job ads, vendor promotion and career content. Roughly a quarter of what looks like consumer conversation is advertising, and filtering it out is a prerequisite for reading demand at all. The full validation funnel.
07 / Influence
Influence is a graph problem, not a prompt problem
Aaru models influence as population level dynamics: conditions applied across a distribution of traits. That is not how a trend moves. Nobody is influenced by an average of a population. People are influenced by the four or five specific trends and events they keep running into that month, and by an object that turns up often enough to stop looking strange. The vector is the trend, not the person carrying it.
So the trend is what we model. A trend is an object plus something happening around it. A plain white t-shirt is one trend when it is worn oversized and half tucked by 19 year olds copying a music video, and a completely different trend when it is bought three at a time and worn under a jacket by someone replacing last year's. Same garment, two audiences, opposite instructions for what to make. Score it as one object and you lose the thing you came for. We score an account by its position in the trend graph, using degree centrality and local network entropy, and by how early and how often it attaches to something that later becomes large. Follower counts are close to useless for this. They measure reach that already happened.
On a graph, propagation is a cascade problem, and cascades have a real literature behind them: independent cascade models, linear threshold models, and graph neural networks that learn the propagation coefficients from observed spread instead of assuming them. That is the architecture, and it is the whole point. The social physics sits in the model, where it can be measured against what actually happened, rather than inside a language model where it can only be imagined.
Nobody decides to start liking a silhouette. They see it forty times in six weeks, it stops looking strange, and the cut they bought last year starts looking dated. That is the mechanism, and no census column records it.
It is also why the rationality problem in section 02 does not come back to bite us. We never ask a model to be irrational on command. We measure what irrational people already did, at scale, and learn how the effect spreads.
08 / What you can run on it
Three decisions, simulated on a real audience
Because the model holds real people, the trends they are attached to and the paths influence travels along, a change can be introduced and propagated rather than merely described.
A product release
Put an unreleased wool overshirt into the map and watch who picks it up first, which communities it reaches in the second and third step, and where it stops moving. The valuable output is not a demand number. It is finding out that the people adopting it first are secondhand-market regulars in Copenhagen rather than the 25 to 34 city commuter the range was drawn for.
A campaign
A campaign is normally judged after it runs, on engagement from the people who already follow the brand, which guarantees a flattering answer. On a graph the question can be asked before the spend: does this leave the first circle, which accounts carry it out, and does it reach anyone who has never bought from you.
A price change
Raising a jacket from 1,200 to 1,600 kronor is not one elasticity across a population. One community stops buying because it is now out of budget. Another starts buying because the price finally matches how they want to be seen. Both movements are visible in the map before the price changes, and they point in opposite directions.
09 / Side by side
The same question, two different machines
| Aaru | Hugo | |
|---|---|---|
| Population comes from | Public and licensed records: census, labour, transactions, visits, search, media use | Observation of real people in public, reached by chain referral through real connections |
| Unit of the model | An agent sampled from a learned joint distribution of traits | A real person, and the trends they are attached to, held on a graph |
| How an individual is represented | A profile drawn from the population structure | A behavioural model learned from what that specific person did |
| Influence between people | Population level dynamics | Cascades over a trend graph, coefficients learned rather than assumed |
| Role of a language model | Not described in detail publicly. Press coverage describes AI respondents | None in the behavioural layer. Used to phrase an answer, never to produce it |
| Demand for something nobody sells | Not represented. No record exists | Visible. People ask for it out loud, and we read it |
| Explains why | By condition and trait, at population level | In the customer's own words, with the post attached |
| Direction in time | Rear facing. Records lag the market | Live. Reads the conversation while it is happening |
| Visual behaviour | Inferred from records | Read directly. Garment level detection from video and images |
| Strongest question | How a documented population splits on a choice that already exists | What to make, launch, say and charge next |
| Interface | Enterprise simulation engagements | Agents that work continuously and report into the brand's own Slack |
| Typical buyer | Corporate strategy, risk, policy, large consultancies | Founders, brand and product teams at consumer brands |
Aaru's method description and the EY figure are taken from Aaru's own public materials and EY's published account of its wealth study recreation.
10 / Which one
A rule that decides it in one question
Ask whether the answer you want is already sitting in somebody's records.
If yes, a record-based simulation is a reasonable instrument. How a documented population splits on an existing choice, what a policy does, how a message lands across a broad demographic spread. The ground truth exists and the job is reconstruction, which is what that method is for.
If no, you need the real population. What people want that nobody makes, why they left you, what they call your category when they are not talking to you, which product should come next and how far it will travel. No amount of statistical sophistication recovers a signal that was never recorded. You have to go where it is being generated, which is people talking and being seen in public.
The fashion and consumer brands we run this for arrive believing they have the first kind of question and almost always have the second. The tell is that the decision in front of them is about something that does not exist yet.
The offer
Bring a question you already know the answer to
The fastest way to judge any of this, ours included, is to ask about an audience you understand deeply and check where the answer disagrees with you. We will run it live on the call and show you the real posts and real people behind every claim.
Related reading: Mapping what consumers actually wear · Building a representative sample · Finding unmet needs · What a consumer simulation is · What social listening misses · Hugo vs traditional market research · All comparisons
Frequently asked questions
Why do AI simulations of consumers fail?
They fail for five structural reasons, none of which a better prompt fixes. A language model is a reasoning engine and consumer behaviour is intuitive rather than reasoned. Its training corpus is dominated by considered written text, which is the least representative sample of consumer behaviour available. Post-training deliberately suppresses the biases that actually move purchases, including status, in-group signalling and gendered assumption. Every agent in the population runs on the same foundation model, so their errors are correlated rather than independent. And training-data leakage makes an honest backtest close to impossible, because almost any study you would test against is already in the weights. The published record agrees: a systematic review across 182 studies concluded synthetic participants do little more than mimic data already collected.
What is the one brain problem in multi-agent simulation?
Running ten thousand AI agents on a single foundation model does not give you ten thousand people. It gives you one brain wearing ten thousand name tags. Every agent draws on the same weights, the same latent space and the same post-training, so their errors all point in the same direction. Averaging correlated agents does not cancel a shared bias, it launders it into a confident number. Research on persona collapse shows this directly: give a thousand agents a thousand distinct profiles and they still converge into a narrow band of behaviour, and pushing fidelity higher makes it worse because the fidelity is bought with stereotype.
Why are LLMs too rational to model consumers?
Language models are trained and then tuned to produce the defensible answer. They evaluate options like an expected-utility calculator and they explain choices with reasons. Real consumers are not doing that. They buy on habit, mood, status, timing, envy, loss aversion and who they saw wearing it first, and then invent the reason afterwards. Alignment training also removes exactly the impulses a researcher needs to see, because a model built to refuse to be prejudiced is a poor instrument for measuring prejudice. The result is a simulation that is right when the answer was already obvious and wrong precisely when the finding would have been worth paying for.
What does Hugo do instead of simulating with LLM agents?
Hugo does not ask a language model to imagine a person. We observe real people in public, reached by chain referral through their real connections rather than by hashtag search, and we read what they actually wear, buy, ask for and complain about, from video and images as well as text. On top of that observed audience we build our own models: a behavioural model per person learned from what that person actually did, and a graph that connects people to the trends they engage with and trends to each other. Influence propagates across that graph as a cascade, learned with graph neural networks rather than assumed. The social physics lives in the model, not in a prompt, so no part of the answer depends on a language model reasoning its way into being a shopper.
Can Hugo simulate a product launch, a campaign or a price change?
Yes, and that is what the influence graph is for. Because the model holds real people, the trends they are attached to and the paths influence travels along, a change can be introduced and propagated: a product release and who adopts it first, a campaign and how far it actually spreads beyond the audience that already follows you, a price change and which segments defect. Aaru answers this class of question by sampling a synthetic population out of records. We answer it by running the change through a map of a real audience and the influence structure between them.
What does Aaru do?
Aaru is a US population simulation company founded in 2024, reportedly valued at a billion dollars, with Accenture as an investor and partner. It builds large synthetic populations of AI agents grounded in public and licensed records, including census and labour data, anonymised transactions, point of interest visits, search demand and media use, then runs scenarios through those populations to predict how a real audience would respond to a price, a product, a campaign or a policy. Its public method description centres on learning the joint relationships between traits and sampling complete profiles from that structure, which is a statistical construction rather than a prompted persona. The response step is described in less detail publicly, and press coverage of its polling work describes thousands of AI respondents answering in under two minutes. EY reported a median Spearman correlation of 0.90 recreating a wealth study across 3,600 affluent investors with it.
Is Hugo an Aaru alternative?
Hugo is the better choice for any consumer brand question about taste, product and culture, which is most of what a consumer brand actually decides. Aaru reconstructs a population from records, so it is strong where the behaviour is already written down: macro policy, well documented demographics, a choice that already exists in transaction data. It cannot contain what nobody recorded, and a product nobody makes yet has no column anywhere. Hugo starts from observation of real people rather than from records, holds a garment-level and language-level map of what they wear, want and reject, and models how influence travels between them. That is the layer where a product launch, a campaign and a price change are actually decided.