Why is it so hard to simulate how consumers behave?
The central problem in simulating consumer behavior is that a language model is trained to be rational, and consumers are not. This is not a flaw in the model; it is its defining feature. A large language model is a reasoning engine. It explains choices with reasons and weighs options like an expected-utility calculator. Real buying decisions, especially in a category like fashion, run on a different operating system: habit, mood, status, envy, and timing. The reason for a purchase is often constructed after the fact. The hard problem, then, is not building an AI that can answer questions, but building an infrastructure that can model the non-rational, social physics of how taste actually forms and spreads. This is the problem we started from.
What are consumer simulations for fashion brands?
Most definitions of a consumer simulation describe the output and skip the part that decides whether it works. The output is easy to state: you ask what if we launched this jacket, what if we raised the price on core denim, what if we seeded this creator instead of that one, and you get an answer before the buy is committed. What separates a useful simulation from a confident guess is the population underneath it, and there are only two ways to get one. You generate it, or you observe it. Generating a synthetic population is fast, and it is what most of this category does. Observing a real one is slow, and it is the only reason the rest of this page can quote an error rate at all. Everything below follows from that single choice.
Why can't I just use a survey or a focus group?
You can, and for some questions, you should. Traditional research methods like surveys, interviews, and focus groups are the incumbent way to test a new idea. But they have three structural problems when applied to the speed of fashion. First, they are slow. A focus group study can take six weeks to return a deck, by which time the trend has moved and the decision window has closed. Second, they are expensive, which means they are commissioned for big bets, not for the dozens of smaller decisions a brand makes every week. Third, and most importantly, they measure what people say they will do, which is often different from what they actually do when no researcher is in the room. A simulation aims to model the behaviour, not the stated opinion.
How does a simulation compare to hiring a research agency?
Hiring a research agency or a panel provider like Kantar or NIQ is the traditional path for major strategic questions. The output is typically a deck, delivered six weeks after the brief. That model has two fundamental gaps. First, the answer arrives after the decision window for most day-to-day questions has already closed. Second, it delivers a document that a human team still has to interpret and turn into a decision. That final step is where the value leaks out. A simulation built on standing infrastructure works differently. It is not a study you commission; it is an agent that runs continuously. It answers a question about a Tuesday decision on Tuesday, and it delivers the recommendation into the place the decision is being made, closing the loop between analysis and action.
Can I use AI agents like ChatGPT to simulate my customers?
The idea is compelling. If a large language model can adopt a persona, it seems logical to create a synthetic panel of thousands of AI-generated customers and ask them what they think of a new collection. You could specify their demographics, their interests, and their past purchases, and get an answer in minutes instead of weeks. This is the promise of synthetic respondents. The problem is that while the answers are fluent and believable, they are often structurally wrong in ways that matter for commercial decisions. The failure is not in the language model, but in the assumption that a reasoning engine can simulate a consumer who is not a reasoning engine.
Why do AI simulations of consumers usually fail?
A language model is trained to be rational. Consumers are not. That is the core of the problem. Real buying decisions, especially in fashion, run on habit, mood, status, envy, and who was seen wearing what first. The reason is often invented after the purchase. When you ask a language model to simulate a shopper, it constructs a rational account of a decision that was never rational to begin with, because constructing rational accounts is the one thing it is designed to do better than anything else. This leads to a simulation that is right when the answer is obvious and wrong exactly when the insight would have been valuable. This is the worst possible error distribution for a research tool. We argue this case in full, against the strongest version of the opposing method and with the published literature cited, in Hugo vs Aaru. If you want the plain explainer first, what a consumer simulation is and how one works covers the ground before this page starts arguing.
Why do synthetic respondents give correlated answers?
A simulation of ten thousand agents built on a single foundation model is not a population of ten thousand minds. It is one brain wearing ten thousand name tags. They share the same underlying weights, the same training data, and the same algorithmic biases. Their errors are correlated, not independent. If the model has a flawed understanding of how a subculture in Brooklyn adopts new styles, all ten thousand of its agents will share that same flawed understanding. Averaging their responses does not cancel out the error; it launders it into a confident-sounding number with a false sense of sample size. The literature measures this as "persona collapse", where distinct profiles converge into a narrow, stereotypical band over time.
Why do AI models miss the social reasons for buying fashion?
Language models are deliberately tuned to be helpful and harmless. This process, called alignment, suppresses their ability to make status judgements, engage in in-group signalling, or make assumptions based on class, gender, or social hierarchy. These are not defects in a shopper; for fashion, they are most of the mechanism. A person buys a certain handbag not just for its utility, but for what it signals to others in their peer group. A model built to refuse to be prejudiced is a poor instrument for measuring prejudice. It cannot simulate the subtle, often unstated social calculations that drive a person to choose one garment over another.
Is a language model trained on the right data to predict fashion trends?
No. The training corpus of a foundation model is dominated by text written by people who stopped to think: Wikipedia, books, news archives, and technical documentation. The closest it gets to raw consumer instinct is a Reddit thread or a product review, and platform licensing changes since 2023 are making that slice of data thinner, not thicker. The real signal for fashion is not in well-formed text. It is in an image, a short video, a comment on a fit pic. As we found in our own research, only 2% of frames in public posts are deliberate "outfit posts". The other 98% is where the real trends live, and that is the part most methods cannot measure. A model trained on the wrong sample cannot predict the behaviour of the right one.
How can you prove an AI simulation is accurate before you use it?
You cannot honestly backtest a large language model's predictions against past events. Almost any published study or historical trend you would use for validation is already in the model's training weights, so a good score measures recall, not genuine prediction. This is a known issue called data contamination. A systematic review of 182 studies on synthetic participants identified four major failure classes: cognitive misalignment, distortion, misleading believability, and contamination. The conclusion was that they do little more than mimic data that was already collected. The only honest way to validate a simulation is against held-out data it has never seen, which is why we publish our own open accuracy log against real, current survey results.
How does Hugo build consumer simulations differently?
We solved the problem by changing the premise. Instead of imagining a synthetic population, we start by observing a real one. We do not generate a persona; we build a small model for each specific, real person in an audience. We do not ask a language model to reason about what someone might do; we simulate the propagation of taste and influence across a graph of real people and their observed connections. The entire process is grounded in observation, not imagination. The social physics of how a trend spreads sits in a measurable model, not inside the black box of a language model.
How do you map an audience without asking them questions?
We build a standing model of a consumer audience from public conversation, images, and video. For our research project, "Mapping what consumers actually wear", we started by reaching real people. We use a method from survey science called chain referral to find authentic members of an audience. For our New York map, this involved starting with a seed set and following real connections to map 876 accounts. We then analyzed their public posts, a total of 4,464 in this case, to understand what they actually wear. This is not a panel; it is an infrastructure holding a live view of an audience.
What kind of fashion data is missed by most research?
In our analysis of public posts from the New York audience, we found that only 2% of image and video frames were deliberate, staged "outfit posts". The other 98% of the data consists of people just living their lives: in the background of a photo, in a candid video, at a party, on the street. This is where the real, un-self-conscious signal of what people are adopting lives. Most methods, human or machine, focus only on the 2% because it is easier to analyze. Our technology is built to extract garment data from the other 98%. From 4,464 posts from 876 accounts, we extracted 9,146 garment crops, clustered them, and found 169 distinct garments that appeared on multiple, unconnected people, showing real-world adoption.
How does a map of real people become a simulation?
The map is the substrate. The simulation runs on top. Each person in the map is a node, and the connections between them are edges representing influence and co-occurrence. When we want to test a new product, we introduce it to the graph. The simulation is not asking a model "what would this person think?". Instead, it runs a cascade problem. Based on the observed behaviour of each person and the structure of the network, which nodes adopt the new product first? How does it spread? Does it jump from one community to another, or does it die out? This is a process grounded in graph theory and social physics, using models like independent cascade and linear threshold, with coefficients learned from previously observed trends.
What kind of fashion decisions can I test with this?
The simulation is pointed at one specific commercial decision. A brand can use it to test a range of questions before committing resources. For example:
- A new product launch: Will this new sneaker design be adopted by the same audience that buys our boots, or will it attract a new customer? Who will be the first adopters?
- A price change: If we increase the price of our core t-shirt by 15%, which communities will leave on budget, and which new communities might arrive because the price now matches their desired status signal?
- A marketing campaign: Will this campaign messaging resonate beyond our core followers, or will it fail to leave the first circle of engagement?
- A creator partnership: Which specific creator will have the most authentic influence in driving adoption of this specific dress among 25-34 year olds in London?
How accurate are these simulations of real audiences?
We continuously validate our simulations against real-world, held-out survey data. In our latest public validation run from June 2026, published in "How accurate are Hugo's simulated polls?", we measured performance on two metrics. On the simple measure of picking the same number one answer as the real survey, Hugo was correct 71% of the time. A more sensitive metric, Jensen-Shannon Divergence (JSD), measures the difference between the full distribution of answers. A score of 0 is a perfect match. Hugo's average JSD was 0.039, indicating that the overall shape of the predicted response was very close to reality, even when the top answer was occasionally missed.
Can this work for audiences outside the United States?
Yes. The method is language and geography-agnostic because it is built on observing real people in their native context. The same validation we run on US audiences, we run globally. For instance, in our open accuracy log, a simulation run on a Finnish audience, seeded with Finnish-language content, achieved a Jensen-Shannon Divergence of 0.028 in a June 2026 run. This score is on par with, and in some cases better than, our results for US audiences. It demonstrates that the approach of mapping real connections and observing behaviour generalizes effectively across different cultures and languages, because the underlying social dynamics are what is being modeled.
What are the honest limitations of this approach?
The method has real, stated limitations. First, it requires public conversation. It cannot work for private populations or for testing a concept so new that nobody is talking about anything adjacent to it. Second, the simulations are sharpest at predicting taste and preference, but they are not a perfect crystal ball. As our own public accuracy report from June 2026 notes, Hugo "tends to pick the right #1 answer but understates how dominant it is, and defaults attitudes toward the middle of the scale." This means it might correctly predict that a black dress will outsell a red one, but underestimate the margin by which it will win. It is a directional instrument for de-risking a decision, not a replacement for final judgment.
How does a simulation result get used in a real decision?
The result is not delivered in a static report or a dashboard that someone has to remember to check. That is the old model, where insight gets lost between the research team and the decision-makers. Hugo agents report their findings directly into a brand's own Slack, into the specific channel and thread where the team is already arguing about the decision. The finding, the recommendation, and the evidence all live in the same place as the conversation. This closes the loop between analysis and action. The goal is not to produce insights, but to give a team a concrete decision, backed by data, in the place they already work.
How can a fashion brand get started with consumer simulations?
Getting started begins with the audience. The first step is for us to build the infrastructure that maps the specific consumer audience your brand sells to. This map of real people, their connections, and what they wear and say becomes the foundation. Once the infrastructure is standing, we can deploy agents on top of it to run simulations against the commercial decisions you have to make, from new products to marketing campaigns to pricing. It is not a tool you log into; it is a capability you point at your hardest problems.