How do AI agents for brands find real customer needs?
An AI agent built for a consumer brand has one job: to help the brand make better decisions about what to make, market, and sell. Most agents fail at this. They provide generic answers because they are built on generic data, or they simulate customers using language models that are trained to be rational in a way real shoppers are not. The hard part is not building the agent. The hard part is building the layer underneath it: a live, standing model of a specific consumer audience, built from what they actually say and do in public. Without that layer, an agent is just a better search engine. With it, an agent can tell a brand what its customers want to buy next, before the sales data catches up.
This is a description of how that infrastructure and the agents on top of it are built. It is based on our own research into mapping real audiences and simulating how trends move through them. Every figure comes from a real run. We explain the method we use, how it differs from prompting a language model to create a synthetic audience, and what its honest limitations are. The goal is to show how an agent can move from answering a question you ask to acting on a decision you need to make.
Why do generic AI assistants give generic answers about customers?
Ask any capable AI assistant what a fashion brand should make next, and you will get a fluent, reasonable, and generic answer. It might suggest "sustainable fabrics" or "vintage-inspired silhouettes". The answer is not wrong, but it is not valuable. It is not specific enough to put in a design brief, and it is the same advice a competitor could get by asking the same question. The problem is not that the AI model is weak. The problem is the data it is standing on.
Most models rely on their general training data, a vast corpus of text and images from books, articles, and the web. That data is a static snapshot of the internet and is not live, meaning it can be months or even a year out of date. Consumer needs are not static. They move daily. A snapshot of the internet from last year cannot answer a question about what a specific audience wants to buy right now. The other option for an agent is to run a live web search, but this limits it to a handful of recent articles and blog posts, which is not a substitute for understanding an entire audience. An agent is only as good as the data infrastructure beneath it. If that infrastructure does not exist, the agent's answers will always be thin.
What is the "one brain" problem with AI personas?
A common workaround for the data problem is to create synthetic respondents. The approach is to prompt a large language model (LLM) to adopt a persona: "You are a 25-year-old living in Brooklyn who loves vintage clothing. What do you think of this new jacket?" By creating thousands of these personas, the thinking goes, you can simulate a customer panel. But this creates a deeper, more subtle problem. Ten thousand agents built on one foundation model is not a population. It is one brain wearing ten thousand name tags.
They all share the same underlying weights, the same latent space, and the same post-training. Their errors are not independent; they are correlated. When one persona misunderstands a cultural nuance, they all misunderstand it in a similar way. Averaging their responses does not cancel out the bias, it launders it into a confident-looking number with a false sense of sample size. An audit of 37 leading models, mentioned in a 2026 preprint, put a number on it: the models resembled each other more closely than any of them resembled the actual human populations they were meant to stand in for. This "persona collapse" means the simulation is most likely to be right when the answer was already guessable, and wrong exactly when the finding would have been worth paying for.
How can you simulate an audience without making them up?
The alternative to imagining an audience is to observe one. Instead of prompting a model into a persona, we build our agents on top of infrastructure that holds a live map of a real consumer audience. The principle is simple: you cannot simulate what you have not measured. We start by mapping the specific audience a brand cares about, resolving it to real people and the relationships between them. This is not a demographic segment named in a workshop. It is a graph of individuals connected by observed interaction and influence.
In one recent project, we mapped a real audience of 40,212 people in Germany to understand how trends spread. We did not ask them questions. We observed their public conversation and behavior to build a mathematical model of their connections and tastes. This model, a graph, becomes the foundation for any simulation. The simulation does not run on imagined personas; it runs on the observed structure of the real world. A trend propagates across the graph according to rules of influence learned from how past trends actually spread within that same group. The output is not what a generic persona might say, but a prediction of what a specific, real person will do, based on their position in the network and their past behavior.
How do you find the real audience for a brand?
Finding a representative group of consumers is a classic research challenge. Search bars and social media analytics provide a view of people who use certain keywords, but this misses the vast majority of an audience. To build a more accurate map, we use a method from survey science called chain referral, or respondent-driven sampling. It is a technique for reaching populations that no public register can list.
We start with a small set of seed accounts known to be part of the target audience. We then follow their public connections and interactions to discover other members of the same community. This process is repeated in waves, building out a network graph of thousands of individuals. For our research mapping what people wear in New York, this method allowed us to map 876 real accounts from 4,464 posts. This is not a panel of people paid to answer questions. It is a map of an organic community, built by following the connections they themselves have made. This ensures the map reflects the real structure of the audience, not an assumed one.
What does it mean to map an audience?
Mapping an audience means building a standing, live model of who they are, what they think, what they need, and what they will buy next. It goes far beyond tracking keywords. It is infrastructure. This map has four layers, each built on the one before.
First, we identify who they are: the real people who make up the audience, connected in a graph. For our New York map, this meant starting with a seed and using chain referral to build a graph of 876 real accounts. Second, we capture what they think, by reading their public conversations about a brand, its category, and its competitors, in the language they use when no researcher is in the room. Third, we identify what they need: the unmet needs and workarounds that signal an opportunity, often visible in complaints or discussions about what a product is missing. Fourth, we determine what they want: the specific product, down to SKU level, that they are ready to buy next. For example, our New York map showed the same specific white sneakers on 7 different accounts, and a particular woven leather bag appearing on 4 accounts in the same cluster. This is demand at SKU level, observed directly from real behavior.
How do you turn photos and videos into data?
A huge portion of consumer life happens in images and video, not text. For a fashion brand, a photo of what someone is actually wearing is a more powerful signal than a tweet about a trend. Yet most research methods ignore this data. Our research into mapping what consumers wear focused on this problem directly. We found that only 2% of public posts containing people were explicit "outfit posts". The other 98% is the unmeasured part of consumer life, and it is where the most valuable signals are.
Our process involves four steps. First, we reach the audience using chain referral. Second, we detect every garment in the photos and videos they post. Third, we use vector similarity search to cluster visually similar garments into distinct items. From 9,146 garment crops in our New York map, we could identify clusters of the same item. Fourth, we resolve these items to real products. This allowed us to identify 832 distinct garments and see how they connected people. For instance, we found 169 garments that were the exact same item, worn by different people across the map, showing real-world adoption patterns that a keyword search could never find.
What is the difference between a trend and just noise?
The internet is full of noise. Thousands of products, ideas, and styles appear and disappear every day. Most are not trends; they are just random fluctuations. An agent that reports every new mention as a trend is worse than useless. The critical task is to distinguish a real, growing signal from statistical noise. To do this, every candidate trend must beat its own baseline.
In our work on audience simulation, we tested 908 candidate trends within a measured audience. For each one, we established a baseline rate of adoption we would expect to see by chance. Only trends that significantly exceeded this baseline were considered real. This statistical trial filtered the 908 candidates down to 578 genuine trends. This rigorous filtering is the first step. It ensures that the agent is acting on real momentum, not just chasing every new thing that appears. Without this step, a brand would be constantly reacting to phantom signals.
How do you measure influence between real people?
Influence is not about follower counts. An account with millions of followers can have less real-world impact on purchasing decisions than a small account that is highly trusted within a specific community. True influence is the measured ability of one person to affect another's behavior. We model this on a graph, where people are nodes and influence is a property of the edges between them.
By observing the sequence of adoptions for hundreds of trends across our audience of 40,212 people, we can learn the coefficients of influence. We see who adopts first and who follows, and we can measure that relationship over thousands of examples. In one case, an account that dwarfed all others by follower count had a low measured influence score. Meanwhile, an account with only 99 followers was the leading influencer across 596 different adoption pairs. The agent learns this from observed data. It does not assume influence based on public metrics. This allows a brand to find the specific creators who can actually move their target audience for a specific product, not just the ones with the biggest audience.
How does this predict what a customer will buy next?
A predictive simulation is the result of putting all these pieces together: the audience graph, the filtered trends, and the learned influence model. To predict what a customer will do next, the agent runs a cascade problem on the graph. It can simulate a product release, a marketing campaign, or a price change, and watch how it propagates through the network of real people.
The models for this come from a real academic literature: independent cascade and linear threshold models, refined with graph neural networks that learn the specific parameters from the audience's past behavior. The agent does not reason about what a "person like this" would probably do. It calculates a predicted reaction for each individual in the graph, based on their own past actions and the influence of their connections. This lets us test scenarios. If a brand launches a new jacket, who adopts it first? Does it spread beyond the initial cluster? Which communities react positively to a new campaign, and which are left cold? The output is a concrete forecast of behavior, not a generic insight.
How does your simulation compare to a language model's?
To test our approach, we ran a direct comparison. We took a real-world dataset where we knew the answer: 7,151 product adoptions that we had held back from our model. We then built two simulations. The first was our graph-based model, trained on the observed behavior of the 40,212 people in the audience. The second was a frontier language model, given the same context and asked to perform the same predictive task.
The results were measured using AUC, a standard metric for predictive accuracy where 0.5 is random chance and 1.0 is perfect prediction. Our fitted model, using the mathematics of the graph, achieved an AUC of 0.86. The frontier LLM, despite being a state-of-the-art model, scored 0.78. More importantly, we found that simply including the network structure of the real audience, the graph of who influences whom, jumped the predictive power for all 7,151 adoptions from 0.58 to 0.81 AUC. This shows that the relationships between people are more predictive of their behavior than their individual attributes alone. The math, fitted to real people, beats the language model trying to reason from a prompt.
What questions can a brand answer with this kind of agent?
An agent built on Hugo's infrastructure of audience maps and simulations can answer specific commercial questions that were previously settled by guesswork. The answers are rooted in the specific technology of observing and modeling real consumer behavior.
- What should we make next? By analyzing the 9,146 garment crops from our New York map, an agent can identify a cluster of visually similar items that people are wearing but no major brand is selling, signaling a specific product opportunity before it has a name.
- Which products in our existing catalogue should we market? The agent can find products a brand already sells that have latent demand, like the woven leather bag that appeared on 4 accounts in our New York map, representing the fastest path to new revenue.
- Which creator or partner will have the most impact? Instead of using follower counts, the agent can use the measured lead score from our simulation model to identify an account with only 99 followers that is the true leading influencer across 596 adoption pairs for a specific audience.
- How will our audience react to a new campaign or price change? The agent can simulate the decision across the graph of 40,212 real people and forecast the reaction with a measured accuracy of 0.86 AUC, showing which communities will adopt it and which will be alienated, before the decision goes live.
How is this different from traditional consumer research?
Traditional consumer research relies on elicited methods like panels, surveys, and focus groups. These are the right tools for certain jobs, like testing a concept for a product that does not exist yet. But they have three structural limits. First, they only answer the question someone thought to ask. They cannot find the opportunity you did not know to look for. Second, they are slow. A study commissioned for a decision on Tuesday often delivers its report six weeks later, long after the decision has been made on instinct. Third, they produce a document. A human still has to read the report, interpret it, and turn it into an action, and this is where most of the value leaks out.
An agent built on standing infrastructure works differently. Because it is always on and observing a whole audience, it can find the opportunity you had not thought to ask about. Because it is live, it can inform the Tuesday decision on Tuesday. And because it is an agent, it does not just deliver a report. It recommends a decision, in the place where the decision is being made, closing the loop between data and action.
How does an agent deliver a decision, not just a report?
Most decisions inside a consumer brand are made on instinct. What to put in the range, what to drop, what to push this month. This is not due to carelessness. It is because the alternative, a six-week study, is too slow for a decision that needs to be made by Friday. The standard output of consumer research, a report or a dashboard, does not solve this problem. It is a document someone still has to read, interpret, and act on, assuming they remember to open it.
An agent closes this loop. It does not live in a separate platform. It reports into the brand's own Slack, into the specific channel where a decision is being discussed. And it does not just provide analysis; it provides a recommended action. It might be a message saying, "Demand for the blue jacket SKU #123 has spiked in this community, based on 14 new posts in the last 48 hours. We recommend increasing the marketing spend on it this week." This is the core difference. It turns consumer intelligence from a strategic project into a daily utility that does work, around the clock, to make sure the brand knows what its customers want.
What are the honest limits of this approach?
This method has clear boundaries, and it is not the right tool for every question. Its primary limitation is that it can only work where people talk and act in public. It cannot measure conversations that happen in private messages or closed groups. This means it is a poor fit for private B2B populations or for categories where discussion is not public.
Second, the signal it reads skews toward the delighted and the annoyed. The frequency of posts is a directional indicator of what people care about, not a statistically representative measure of incidence across a whole population. For a regulated claim that requires a defensible sample, a traditional panel or survey is a better instrument. Finally, it cannot test a true counterfactual for something with no signal at all. To test a concept for a completely new product that does not exist in any form, elicited research, asking people directly, is a more suitable method. Our agents are built to observe and simulate the world as it is, not to test ideas that have no observable footprint.
How do you know any of this is real?
Believability is a major failure mode for AI-driven research. Synthetic respondents produce fluent, believable answers that can be completely wrong. A core principle of our approach is auditability. Every claim an agent makes should be traceable back to its source. When an agent identifies a trend or a piece of product feedback, it does not just provide a summary. It provides a direct link to the real, public post where the observation was made, complete with the account and the date.
A brand can click through and see the evidence for themselves. This is the strongest filter in the category. It prevents the agent from "hallucinating" findings or laundering a model's bias into a confident assertion. If a claim cannot be backed by a link to a real person saying or doing something in public, it does not get surfaced. This makes the entire system accountable to reality and ensures that the brand is making decisions based on real human signal, not a model's invention.