Why is simulating real consumer behavior so hard?
The core challenge of consumer simulation is that a language model is trained to be rational, but consumers are not. The standard approach prompts a large language model to adopt a persona, but this generates a plausible story, not a predictive model of irrational human choice. A 2026 audit of 37 foundation models found they resembled each other more closely than any of them resembled the real human populations they were meant to simulate. The hard part is not generating a fluent answer; it is building a model that captures the non-rational, networked behavior that actually drives a purchase. Our approach starts by observing it: for one German audience, we mapped 40,212 real people to build a mathematical model of influence, not a fictional persona.
How do most AI tools try to simulate consumers?
The standard approach today is to use a large language model (LLM) like those that power ChatGPT or Claude. The method is straightforward: you write a prompt that describes a target customer in detail. You might specify their age, income, location, hobbies, and brand preferences. You then ask the LLM to adopt this "persona" and answer questions or react to scenarios as that person would. To create a simulated audience, you simply repeat this process, generating hundreds or thousands of these AI personas to stand in for your real customer base.
Why do LLM personas seem so believable?
Language models are masters of fluency. When you ask one to role-play as a 28-year-old rock climber from Denver who buys sustainable clothing, it produces a response that sounds exactly like you would expect. The language is correct, the references are plausible, and the reasoning is articulate. This creates what researchers call "misleading believability". The output is so coherent and well-written that it feels true, even when it has no connection to the messy, irrational reality of how people actually think and buy. The polish of the answer is mistaken for the accuracy of the simulation.
What is the core failure of LLM persona populations?
The fundamental problem is this: a language model is trained to be rational, but consumers are not. An LLM is a reasoning engine. It constructs logical arguments and explains choices with neat, ordered reasons. Real human buying decisions are driven by habit, mood, status anxiety, envy, what a friend wore, or a dozen other factors that have little to do with a rational calculation of utility. The reason is often invented after the purchase, not before. When you ask a model to explain a buying choice, it defaults to what it does best: it constructs a rational story for a decision that was never rational to begin with. This is the shallow half of the problem. The deeper, structural half is what we call "one brain".
Why is having 'one brain' a problem for simulation?
Ten thousand AI agents built on a single foundation model are not a population. It is one brain wearing ten thousand name tags. They all share the same underlying weights, the same training data, and the same post-training adjustments. This means their errors are not independent; they are correlated. When one agent gets something wrong for a structural reason, they all get it wrong in the same way. Averaging their responses does not cancel out the noise to reveal a signal. It launders a shared, systematic bias into a confident-looking number, creating a false sense of security from a large sample size. A 2026 audit of 37 models from labs like OpenAI, Google, and Anthropic found that the models resembled each other far more closely than any of them resembled the real human populations they were meant to simulate.
How does AI 'alignment' break consumer simulation?
Language models are deliberately tuned to be helpful and harmless. This process, often called alignment or RLHF, actively suppresses many of the core mechanisms of consumer behavior. The models are trained to avoid making status judgments, using stereotypes, showing prejudice, or making class-based inferences. But these are not bugs in a shopper. For better or worse, they are a significant part of the mechanism. A desire to signal belonging to an in-group, a judgment about what a certain product says about its owner, or a gendered assumption about a category are powerful drivers of real-world purchasing. A model that is explicitly built to refuse to think in these ways is a poor instrument for measuring their effect.
Is an LLM trained on the right data to simulate a shopper?
No. The training corpus of a large language model is overwhelmingly composed of text written by people who stopped to think: Wikipedia articles, books, news archives, technical documentation, and structured arguments. It is a library of considered thought. The closest the data gets to raw consumer instinct is in forum threads or product reviews, but even that is a sliver of the whole. The unfiltered, impulsive, low-context chatter of a TikTok comment section or an Instagram story is not the dominant signal. Since 2023, platform licensing has made this slice of real-time, un-edited consumer conversation thinner, not thicker, for model trainers. The model is learning from the wrong sample of human expression.
Can you honestly backtest an LLM-based simulation?
It is exceptionally difficult. The problem is data contamination. Almost every major published consumer study, every market research report, and every academic paper from the last few decades is likely already in the model's training weights. If you try to validate the simulation by asking it to recreate the results of a known study, you are not measuring its predictive power. You are measuring its ability to recall information it has already memorized. A good score reflects recall, not accurate prediction. Any vendor claiming clean historical validation of an LLM population should be asked to prove how they ruled out this kind of data leakage. Without that proof, the validation is meaningless.
What is the practical result of these failures?
The combination of these flaws leads to a predictable and dangerous pattern of error. The model produces the most conventional, stereotyped answer for the persona described. This means the simulation is often correct when the answer was already obvious and easily guessable. It fails in exactly the moments where the insight would have been non-obvious and valuable enough to pay for. For a brand trying to find an edge, this is the worst possible error distribution. It confirms what you already know and misleads you when you are entering unknown territory.
If personas fail, what is the alternative to LLM simulation?
The alternative is to ground simulation in observation, not imagination. Instead of prompting a model to pretend to be a customer, we start with real customers and real behaviors. Our principle is simple: you cannot simulate what you have not first measured. This means building a model from the ground up, based on the observed actions of a specific audience. It is the difference between writing a fictional character and building a statistical model of a real person. The output is not a fluent paragraph of prose; it is a predicted reaction, a probability, a score. It is math, not literature.
How can you observe an audience without asking them questions?
We read what people choose to share in public. Our infrastructure is built to understand consumer conversation, but also images and video. In our work mapping fashion audiences, we found that only 2% of relevant posts were explicit "outfit posts". As we reported in our study, Mapping what every consumer audience actually wears, the other 98% is where the real signal is: people just living their lives, incidentally wearing the clothes they actually own and prefer. We built technology to detect garments in ordinary photos, cluster them into distinct items, and map how they are connected across thousands of people. This creates a detailed, evidence-based picture of what an audience wants before it ever shows up in sales data.
How do you find a representative audience to observe?
You cannot find a real audience by typing keywords into a search bar. Instead, we use a method from survey science called chain referral, or respondent-driven sampling. We start with a small, verified set of people who fit the audience definition, and then follow the real social connections outward. This allows us to reach the clusters and communities that make up an audience, rather than just the loudest individuals. For our New York map, this process involved analyzing 4,464 posts from 876 accounts to build a verifiable picture of what that specific community wears, thinks, and wants.
How does Hugo model influence and trends?
We build a graph. In our model, the nodes are real people and emerging trends, and the edges represent influence and adoption, measured from real data. As we detailed in our research on audience simulations built on math, we mapped a German audience of 40,212 real people and tracked how 578 distinct trends propagated through the network. We do not create a generic persona. We build a small, specific statistical model for each person, trained only on what that individual has posted, worn, and reacted to. Influence is not a guess; it is a coefficient in a model, learned from observed data about who adopts what after whom. The social physics sits in the mathematics, where it can be measured, instead of inside a language model where it can only be imagined.
Can you prove this mathematical approach is better?
Yes. We test our methods on a world where we already know the answer. We take a set of real trends that spread through a real audience, and we hold back the most recent adoptions. The test is to predict who will adopt the trend next. In our published research, we gave a frontier language model the exact same historical context and persona information that our mathematical model received. Our fitted model predicted the correct adopter with an AUC score of 0.86. The LLM scored 0.78. More importantly, when we switched on the network information from our graph, the part that understands who influences whom, the accuracy of our model jumped from a baseline of 0.58 to 0.81 across all 7,151 held-out adoptions. The numbers show that observing the real social structure is not just an alternative; it is measurably better.
What kind of questions can this type of simulation answer?
Because the simulation is built on a graph of real people and their connections, it can answer questions about propagation and cascades. For a brand, this means running experiments that are impossible with static personas. You can simulate a product release and see which communities adopt it first and who the key individuals are that spread it to the next circle. You can test a marketing campaign and see if the message is likely to stay within an initial echo chamber or break out into the wider audience. You can model a price change and see which customer segments are most likely to leave, and which might actually be attracted by the new positioning.
How is this different from other AI consumer research tools?
Most AI consumer research tools are designed to answer a question you ask. They might analyze survey responses, track keyword mentions, or summarize product reviews. The output is typically a report or a dashboard that you have to interpret and then decide how to act on. Hugo is different. The simulation layer is part of a standing infrastructure that is always on. Agents built on top of this infrastructure do not wait to be asked. They monitor the audience around the clock and surface findings proactively. The deliverable is not a report; it is a decision recommendation, sent directly into a brand's Slack, into the thread where the team is already discussing the issue. It closes the loop from insight to action.
How can you be sure the influence you measure is real?
Real influence is often invisible to simple metrics like follower counts. In our German audience simulation, the account with the most followers was not the most influential. Our model, which measures influence by tracking who actually causes others to adopt trends, identified a different account with only 99 followers as the critical node for spreading 596 different adoption pairs. This is the kind of non-obvious finding that can only come from measuring the network structure, not from assuming that reach equals influence. The simulation works because it is based on who people actually listen to, not who has the biggest megaphone.
What are the limitations of observing real people?
This approach has clear and important limits. It only works for topics that people discuss and behaviors they exhibit in public. It cannot read minds, and it cannot measure the opinions of people who do not post. The data naturally skews toward the delighted and the annoyed, so while it is directionally powerful for understanding "what" and "why," it cannot be used to definitively state "how many." It is an instrument for understanding the dynamics of demand, not for producing a census-representative statistic. Finally, it cannot test a true counterfactual for an idea that has no precedent in the market, because it is grounded in observing what is, not imagining what might be.
How can I see the evidence for myself?
Auditability is core to our approach. We believe a brand should be able to trace any claim back to its source. In our New York wear map, every finding is connected to its evidence. When we show that a particular pair of white sneakers is a key item for an audience, you can see the seven accounts that wear it. Clicking on a garment opens the cluster of all 9,146 garment crops we analyzed, allowing you to see the vector search results that our human reviewers checked. The portraits are illustrations, but they link to the real public accounts. We provide the receipts. It is the opposite of a black box.
How can my brand use this kind of simulation?
The purpose of our infrastructure is to give your brand the ability to test its most important decisions against a live, accurate model of your specific audience. It replaces boardroom debates and gut feelings with testable hypotheses and data-driven answers. We build the graph of your audience, we model the dynamics of how taste and information spread, and we give you the tools to run the experiments. If you are a consumer brand that needs to know what to make next, how to market it, and who will move the needle, our approach is designed for you. Bring an audience. We will show you who moves it.