Blog / Consumer intelligence

How to simulate a fashion audience: the flaws of AI and the power of observation

This article explains why large language models are structurally flawed for simulating fashion audiences. It outlines an alternative approach grounded in real-world observation and mathematical modeling of consumer behavior.

By the Hugo team · Published 29 August 2026 · 14 min read

How do you simulate a fashion audience?

To simulate a fashion audience, you first have to measure it. A simulation is only as good as the data it stands on, and for fashion, that data cannot be a series of rational statements. Real buying decisions run on habit, mood, status, and timing, with the reason invented afterwards. Most attempts at simulation begin by prompting a large language model to act like a consumer, but this approach is structurally flawed. A language model is a reasoning engine trained to be rational. Consumers are not.

The alternative is to observe what a real audience wears, says, and reacts to, and then build a mathematical model of their behavior. This means mapping real people and the garments they wear from public conversation, images, and video. It means building a graph of who influences whom. Only then can you run a simulation that predicts what will spread, who will adopt it, and why. This piece explains how that works, using real numbers from our own published research.

Why can't I just ask an AI assistant to simulate my customers?

You can, and the answer will be fluent, generic, and wrong in the one way that matters. An AI assistant like ChatGPT or Claude, when asked to simulate a consumer, is performing a kind of role play. It constructs a persona based on the stereotypes in its training data and generates a plausible answer. The problem is that this process misses the entire mechanism of real consumer choice, especially in a category like fashion that is driven by non-rational factors.

A language model is an expected-utility calculator. It weighs options and explains its choices with reasons. But a real shopper does not buy a jacket after a rational deliberation of its features. They buy it because of who they saw wearing it, how it made them feel, or what it signals about their identity. When you ask a model to explain that purchase, it will invent a rational account of a decision that was never rational. This is because constructing rational accounts is the one thing it is built to do. The result is an answer that is right when the outcome was already obvious, and wrong exactly when the insight would have been valuable.

Are simulated consumer populations really diverse?

The "one brain" problem is the critical flaw in using language models to create a synthetic population. Ten thousand simulated agents built on a single foundation model are not a population of ten thousand minds. They are one brain wearing ten thousand name tags. They all share the same underlying weights, the same latent space, and the same post-training alignment. Their errors are not independent; they are correlated.

In real research, averaging responses across a large sample helps cancel out individual biases and noise. But with a synthetic population from a single model, averaging does not cancel out a shared bias. It launders it into a confident-sounding number, creating a false sense of security from a large sample size. A 2026 audit of 37 different language models found that the models resembled each other more closely than any of them resembled the actual humans they were supposed to be simulating. This convergence, sometimes called persona collapse, means you are not measuring a diverse audience. You are measuring the internal consistency of a single machine.

How does AI alignment make consumer simulation worse?

The very process used to make language models safe and helpful actively deletes the mechanisms that drive consumer behavior. Models are deliberately tuned during alignment to suppress status judgment, in-group signaling, gendered assumptions, and class inference. The model is taught not to be prejudiced or make snap judgments based on appearances.

But those are not defects in a shopper. They are most of the mechanism. Fashion is a language of status, identity, and belonging. A model that refuses to engage in that language is a poor instrument for measuring it. A recent study in a 2026 preprint found that across 70,000 real survey answers, prompting a model with a detailed persona often made its answers a worse match for real people, because the alignment training overrode the persona. The model's core instruction to be neutral and helpful fought against the instruction to be a specific, biased human. For a brand trying to understand the messy, real-world drivers of desire, a perfectly aligned model is the wrong tool for the job.

Is the training data for language models a good sample of consumers?

No. The corpus a language model is trained on is overwhelmingly made of text written by people who stopped to think. It is Wikipedia, news archives, books, scientific papers, and technical documentation. It is edited, structured, and rational. The closest it gets to unfiltered consumer instinct might be a Reddit thread or a product review, but even that is a sliver of the whole.

Since 2023, platform licensing has made that slice of real, spontaneous conversation thinner, not thicker. The model's diet is increasingly made of its own output and other formal text. It is not trained on the fleeting visual signals of a TikTok video, the nuance of an outfit in a photo, or the specific language people use when they are trying to find a product that does not exist yet. The data is the wrong sample, which means the model has a fundamentally skewed view of the world it is asked to simulate.

What does it mean to observe an audience instead of imagining one?

Observing an audience means starting with real people, not a prompt. It means your simulation is grounded in measured reality, not a work of fiction. Instead of asking a model to imagine what a 25-year-old in London might wear, you go and find out. At Hugo, this is the foundation of our entire approach.

Our work begins by mapping who consumers are, what they think, what they need, and what they will buy next, at a product level. This infrastructure is built from public conversation, images, and video. For example, our New York map processed 4,464 posts from 876 accounts to identify 832 distinct garments people were actually wearing. It is not a study you commission; it is an always-on map of your audience. The simulation runs on top of this map, not on a generic language model. This is the difference between astronomy and astrology. One is a model built from observation; the other is a story told about an imaginary system.

How do you map what a fashion audience actually wears?

We map what people wear by looking at them. It sounds simple, but it is a complex data science problem that no one else has solved at scale. For any audience a brand cares about, we build a map in four layers. We published our work on this in a piece called "Mapping what every consumer audience actually wears". For our New York map, we started with real people. We used a survey science method called chain referral to reach 876 real accounts, not just influencers. From their 4,464 posts, we analyzed every frame.

Crucially, we found that only 2% of frames were explicit "outfit posts". The other 98% is the ordinary life where real trends live, and it is the part that every other tool misses. Our models detected and cropped 9,146 garments from these images. We then used vector search to cluster these crops, identifying 832 distinct garments and finding 169 of them being worn as the exact same item across multiple, unconnected accounts. This creates a graph not of what people say they wear, but of what they are actually seen in, connected to the real people wearing it.

What kind of data comes from mapping a real audience?

Mapping a real audience produces a graph of people, products, and the relationships between them. It is not a dashboard of sentiment scores or keyword trends. It is a detailed, structured model of a market. From our New York map, you can see that the same white sneakers appear on seven different accounts. You can see that the people wearing those sneakers are also wearing a specific crew tee, a style of jeans, and a particular brand of leggings. You can follow the connections to see a woven leather bag appear on four other accounts in the same cluster.

This is not an abstract segment called "sneaker wearers". It is a specific group of real people, connected by shared taste, observed in the wild. You can see the exact garments, click through to the evidence, and understand the context. This infrastructure tells you what products are being worn together, which items define an aesthetic, and who is central to that aesthetic. It is the ground truth needed to simulate how a new product or trend might move through that same network.

How can you simulate a trend before it happens?

You simulate a trend by modeling the mechanics of influence and adoption on a graph of real people. A trend is not a disembodied idea; it is a piece of information or a behavior that propagates through a network. Our approach, which we detailed in our research on audience simulations, is to model this propagation mathematically.

First, we identify candidate trends from the observed data. In a study of one German audience of 40,212 people, we tested 908 candidate trends and found 578 that survived statistical trials, meaning they were spreading non-randomly. Then, we model two things: influence and adoption. For every pair of people in the network, we calculate an influence score based on observed interactions. For every person, we calculate an adoption threshold based on their past behavior. A simulation then becomes a cascade problem: a trend "infects" the early adopters, and we watch to see if it spreads across the network, crossing the thresholds of their neighbors. This lets us predict which of the 578 trends will take off and which will die out.

How do you measure influence in a fashion audience?

Influence is not the same as follower count. The most influential person in an audience is the one who most often precedes others in adopting a new trend. We measure this directly from the data. In our German audience study, we looked at every adoption of every trend and traced it back to who in that person's network had adopted it first. This creates a "lead score" for every person in the graph.

The results are often surprising. In that audience, one account dwarfed all others by follower count. But when we measured influence directly, that account's importance shrank. A different account, with only 99 followers, was the leading indicator for 596 different adoption pairs across the network. They were a true innovator, whose choices were predictive of what others would do next. A simulation built on follower counts would get this wrong. A simulation built on a measured graph of influence gets it right. This is the difference between measuring vanity metrics and measuring the actual physics of a market.

Can you prove this approach is more accurate than an AI assistant?

Yes. We test every method on a world where we already know the answer. For our simulation research, we used a real, historical dataset of 40,212 people and 7,151 trend adoptions. We held back a portion of this data and tasked our models with predicting who would adopt which trend next. We then compared our results to the ground truth of what actually happened.

Our fitted mathematical model achieved an Area Under the Curve (AUC) score of 0.86 in predicting adoptions. An AUC of 0.5 is random chance, and 1.0 is a perfect prediction. We then ran the same test with a frontier large language model. We gave the LLM all the same context our model had: the person's profile, their past behavior, the trend itself. The language model scored 0.78. Our model, built on the observed social physics of the network, was significantly more accurate. In another test, we measured the uplift from including the network information. Without it, our model scored 0.58. By switching on the measured graph of influence, the score jumped to 0.81. This proves that who you know is a better predictor of what you will do than who you are.

What questions can a fashion audience simulation answer?

A simulation built on a real audience map can answer the concrete commercial questions a brand faces every day. It is not for generating abstract "insights". It is for making decisions. The entire infrastructure exists so a brand can ask four things:

  1. Product Release: If we launch this new jacket, who adopts it first, and does it spread beyond them? You can test different designs, colorways, or price points and see how they propagate through the real audience graph.
  2. Campaign Impact: If we partner with this specific creator, will their message leave their first circle of followers? You can simulate the ripple effect of a campaign and see if it reaches the communities you actually want to sell to.
  3. Price Change: If we raise the price of our best-selling dress, which communities will leave on budget, and which new communities might arrive because the price now matches how they want to be seen?
  4. Competitor Threat: A rival brand just launched a new sneaker. Our simulation can show you how quickly it is spreading through your own customer base and who the key carriers are.

These are not questions about what people are saying. They are questions about what to do next, answered with a defensible, quantitative forecast.

How is this different from AI consumer research tools?

Most AI consumer research tools fall into one of two camps: elicited or observed. Elicited tools, like AI-moderated interviews, are designed to ask people questions about things that do not exist yet, like a new concept or logo. They help you test a hypothesis you already have. Observed tools, like social listening platforms, are good at measuring keywords you can already name. They can tell you the sentiment around your brand name or the share of voice for a campaign hashtag.

A simulation is different in kind. It does not just report on the past or ask about a hypothetical future. It models the causal mechanics of a market to predict the likely future. It answers the question "what happens if we do this?". While other tools give you a report or a dashboard that you then have to interpret and act on, a simulation gives you the likely outcome of a decision before you commit to it. It delivers a forecast for a specific decision, helping a team choose what to do next.

How is this different from trend forecasting services?

Traditional trend forecasting services, like WGSN, operate at a macro level. They publish seasonal reports on broad trends, colors, and silhouettes that every subscriber reads at the same time. This is useful for context, but it is not specific to your brand, your audience, or your product. It tells you that "dopamine dressing" is a trend; it does not tell you if your customers will buy a bright yellow version of your handbag.

A simulation on a real audience map works at the micro level. It is specific to your audience and your catalogue. It can tell you not just that a trend exists, but whether it has traction with the specific people who buy from you. It can identify a trend that has no name yet, emerging from the co-occurrence of specific garments within a small but influential cluster of people. It moves the work from consuming a generic forecast to testing a specific action against your specific market.

What are the limitations of this kind of simulation?

The honesty of any model depends on stating what it cannot do. Our approach has clear boundaries. First, it only works where people talk and show their lives in public. For private B2B populations or decisions made behind closed doors, the method has no signal to read. Second, the data skews toward the delighted and the annoyed. The frequency of posts is a directional signal of what people care about, not a representative measure of incidence in a whole population. We can tell you what is gaining momentum, but we cannot tell you exactly how many people have bought it.

Finally, a simulation can test a decision, but it cannot test a true counterfactual. We can simulate what happens if you launch a blue jacket, but we cannot simulate what would have happened if the color blue had never been invented. The model is grounded in the observed world and can only extrapolate from it. It is a tool for navigating the future, not for rewriting the past. We believe being plain about these limits is what makes the results trustworthy.

How can I get a simulation for my own fashion audience?

The first step is to map the audience. A simulation requires a substrate of real people and their connections to run on. We work with brands to define the audience they care about, whether it is their existing customers, a competitor's, or a demographic they want to win. We then build the map for that specific audience, following the same process of chain referral, garment detection, and network construction that we have detailed in our research.

Once that infrastructure is in place, it is permanent. It is a standing model of your market that updates as the audience changes. On top of that, we build and deploy agents that run the simulations you need. An agent is the infrastructure pointed at one commercial decision, like a new product launch or a marketing campaign. It lives in your brand's own Slack, reporting its findings where your team already makes its decisions. It is not another dashboard to check. It is an active participant in your daily work, providing a quantitative check on the instincts that drive your brand forward.

Bring the question you are actually stuck on.

We will tell you honestly whether it is an observed-signal question. If it is, we will run it live on the call and you can judge the answer.