Compare / Simile
Hugo vs Simile
Hugo and Simile are both in the business of simulating consumers and predicting what they will do. Simile builds a digital twin of a person from a long interview and lets a company put questions to millions of them. We watch what real people wear, buy and ask for in public, fit a behavioural model to each of them, and simulate how a product or a campaign travels between them. Both companies start from real people, which most of this category does not. We part ways on what a prediction should be built from, and on what happens once the simulation has made one.
Published 19 September 2026. Our account of Simile's method comes from their public site, their founders' published research and press coverage, so read it as a careful reading and not an audit.
Short answer
Choose Simile if the thing you need to predict is an attitude or a stated choice, across a population you can recruit and interview: how members respond to a policy, which message persuades a patient, how a bank's customers talk about switching. Their founding research is the most careful work published on simulating individuals, and they validate against real people every week.
Choose Hugo if you run a consumer brand and the decision is what to make, what to launch, who to put it on and what to charge. Those outcomes are set by what people do and by who they copy, and neither of those comes out of an interview. We also do not stop at the answer. Agents built on the same infrastructure work the decision every day and tell your team when something changes.
The difference in one line: Simile simulates what people would say. We simulate what people do, and then we act on it.
01 / Common ground
Both of us start from real people, and that already separates us from most of the field
The cheap version of consumer simulation is a language model given a paragraph describing a shopper and asked to answer as her. Nobody real is anywhere in it. Simile does not do that, and neither do we.
Simile's populations begin with recruited humans who sat through a real interview. Ours begin with real accounts reached through their actual connections. In both cases there is a specific person underneath every simulated one, which means the simulation can be checked against that person. That is the entry ticket for taking any of this seriously, and it is why this page argues with Simile's method in detail. Aaru builds its population a third way, from records, and we have written about that separately.
02 / What Simile actually does
Their method is the best version of asking
Simile was founded by Joon Sung Park, whose 2023 Stanford paper on generative agents put 25 language model characters in a small simulated town and watched them form routines, relationships and a party nobody scripted. It is one of the most cited papers in the field, and the company is the commercial continuation of that research line.
The work that matters for buyers came a year later. In a 2024 study, Park and colleagues interviewed 1,052 Americans for two hours each, built a language model agent from every transcript, and then tested whether each agent answered the General Social Survey the way its human did. The agents reached roughly 85 percent of the accuracy the participants themselves managed when they retook the same survey two weeks later.
That denominator deserves credit. People do not agree with their own answers from a fortnight ago, so scoring an agent against a perfect record would be measuring against something no human achieves. Normalising by test and retest is the honest way to report it, and most vendors in this category do not bother.
The company built on that has raised a reported 100 million dollars led by Index Ventures and then 200 million at a two billion dollar valuation led by Greenoaks, within about six months of each other. Its site names CVS Health, Gallup, Deloitte, Wealthfront, Itaú and Suntory as customers, and press coverage describes CVS running 400,000 twins to study how to get patients to take their medication. They describe validating against real humans weekly and tagging each result with a predicted accuracy level.
If the behaviour you need to predict is an answer to a question, this is a serious instrument, and we would not try to talk you out of it.
03 / The first disagreement
An interview measures what people say about themselves
The 85 percent figure compares two self-reports: what a person told a survey, and what their twin told the same survey. It shows the twin has learned how that person describes themselves. It does not show the twin has learned what that person buys, because at no point in the benchmark does anybody buy anything.
For a lot of decisions that gap does not matter. For a consumer brand it is most of the problem, and the research on it is old and settled.
- People cannot report their own reasons. Nisbett and Wilson's 1977 review, Telling more than we can know, showed across dozens of experiments that people confidently explain choices using causes that demonstrably did not operate, and miss the ones that did. Two hours of careful interviewing collects those explanations at length.
- Intention is a weak predictor of action. Sheeran's 2002 review of the intention and behaviour literature found stated intentions account for roughly 28 percent of the variance in what people then do. The rest is habit, circumstance and everything the person did not think to mention.
- Clothing is close to the worst case. Ask someone why they bought the jacket and you get fit, fabric and price. You will not hear that they had seen that cut forty times in six weeks on people they would like to resemble, because they do not know that themselves. Status and imitation are not things people withhold in an interview. They are things people have no access to.
So a twin built from an interview inherits the interview's blind spot with great fidelity. It will tell you, in that person's voice, the same story that person would have told you. For a brand, that story is the part of the customer you could already reach with a survey.
We start from the other side. In our New York map, 2 percent of the frames we read were posts about clothing. The other 98 percent were people filming a kitchen, a dog or a night out while wearing something. Nobody chose those clothes for a researcher, which is exactly why they are worth reading. The full run is published.
04 / The second disagreement
A panel is a room full of strangers, and taste moves between friends
A well built panel is recruited so that its members are independent of each other. That is what makes it representative, and it is the right design for measuring opinion. It also removes the one mechanism that decides whether a product spreads.
Nobody adopts a silhouette alone. They adopt it because specific people upstream of them did, and whether your launch travels depends on who those people are and who is downstream of them in turn. A million independent twins can each give a considered view of your overshirt, and the sum of those views is a tally of opinions. It has no way of expressing that the eleven people who would wear it first are all watched by the same four thousand.
That structure is the object we model. The German audience in our published simulation work is 40,212 real people joined by 60,015 confirmed follow edges, and the prediction task is stated on that graph: which trends spread, who adopts, in what order. Influence is fitted from spread that already happened and then scored against a future the model never saw. We covered the cascade mathematics on the Aaru page and will not repeat it here.
05 / The third disagreement
The twin still answers through a language model
An interview transcript gives a language model far better material than a persona paragraph does. It is still a language model producing the response, and that matters most for the class of decision a consumer brand makes.
A language model is trained and then tuned again to give the defensible answer. It is very good at reasons. Buying clothes runs on mood, timing, envy and who was seen in it first, with the reasons assembled afterwards. Give that model a transcript full of a person's own after-the-fact reasons and you have stacked one rationalisation on another. The output will be plausible and articulate, and it will lean toward the sensible answer in exactly the cases where the real person would not have been sensible. The five structural reasons, with the published evidence, are on the Aaru page.
Simile's site says its agents take in news and media as the people they represent would, which is their answer to the obvious objection that an interview goes stale. It is a fair answer for opinions about public events. For taste, the update that matters is not what the person read this week. It is what the people around them started wearing, and a twin has no people around it.
There is no language model in our behavioural layer. Each person carries a small model fitted to what that person actually posted, wore and reacted to, and it outputs a predicted reaction directly. A language model is used at the end to phrase a finding for a human reader, never to produce it.
06 / After the answer
A prediction is only worth what acts on it
Set the method aside entirely and there is a second difference, which for most brands is the larger one. Simile and Aaru both sell a simulation you query. Someone has to think of the question, run it, read the prediction, and carry it into the meeting where the decision is made. The simulation can be excellent and the launch still goes out on instinct, because the prediction arrived in a different room from the decision.
We treat consumer simulation and prediction as infrastructure and build agents on top of it. An agent is that infrastructure pointed at one commercial decision for one brand, running continuously, reporting into the brand's own Slack in the thread where the decision is being argued. Our engineers work inside the brand and build each agent for what that brand needs. The four below are among the ones brands ask for first.
What to make next
Reads what your audience is asking for and cannot find, and recommends a specific item you could brief a factory on.
What to fix in the current range
Feedback on what you already sell, in the customer's own words, including the complaints that never reached your support inbox.
What to push this week
Finds the product already in your catalogue that your audience has started wanting. The fastest revenue in most catalogues is something the brand already makes.
Who should wear it
Identifies the one account that would move one product for one audience, scored by position in the graph rather than follower count.
Every agent is custom, because a decision is specific. There is also a simulation agent, which tests a concept, a price or a campaign on your audience before you spend, and whatever your team needs built. The agents page has the detail and the error rates underneath them.
This is why we spend most of our effort on the prediction layer. Building an agent is easy now, so two agents differ only in what they know about your customers. Ours know what your audience wore yesterday and who they copied it from.
07 / Side by side
The same customer, reached three different ways
| Simile | Aaru | Hugo | |
|---|---|---|---|
| What it predicts | What people would say if you asked them | How a documented population splits on an existing choice | What people will wear and buy next, and how it spreads between them |
| Population comes from | Recruited people, interviewed at length, extended with panel data | Public and licensed records: census, transactions, visits, search | Real people observed in public, reached by chain referral through their connections |
| What is known about each person | What they said about themselves | A profile sampled from population structure | What they posted, wore and reacted to |
| What produces the response | A language model grounded in the interview | Not described in detail publicly | A behavioural model per person. No language model in that layer |
| Links between people | None. A panel is independent by design | Population level dynamics | A follow graph, with influence fitted from observed spread |
| Published validation | Agreement with the person's own survey answers, normalised by test and retest | Recreation of a published wealth study | Held-out surveys, and adoption scored against a future the model never saw |
| Strongest question | How will these people answer, and which message moves them | How a documented population splits on an existing choice | What to make, launch and charge, and who carries it |
| What you receive | Predictions from a simulation you run | Predictions from a simulation engagement | Agents that work continuously and report into your Slack |
| Typical buyer | Health, financial services, consultancies, insights teams | Corporate strategy, agencies, policy | Founders, brand and product teams at consumer brands |
Simile's method, customers and funding are taken from Simile's public site, its founders' published papers and press coverage. Aaru's row is sourced on its own page.
08 / Where we lose
What Simile can do that we cannot
We only see people who are visible in public, and no roadmap item changes that.
- Private behaviour. Whether a patient takes a pill, how someone chooses a pension, what an employee thinks of a reorganisation. None of that is posted anywhere. A recruited panel reaches it and we do not.
- A defensible representative sample. If a regulator or a board needs a sample with a documented recruitment frame, a panel has one. Our audiences are reached by chain referral, which is a published survey science method, and the degree-weighted bias correction that method calls for is designed and not yet implemented in our pipeline.
- Audiences that do not post. Public observation over-represents the young, the delighted and the annoyed. Frequency in our data is directional, and it should not be read as incidence.
- Your own first-party data. Simile accepts opt-in loyalty and account data under the customer's governance. We work from public signal and do not ingest a customer file.
If your question lives in any of those four, call them first.
09 / Which one
A rule that decides it in one question
Ask whether the outcome you care about is something a person could tell you.
If they could, simulate the asking. Attitudes, stated preferences, reactions to a message, a policy or a plan. A twin built from an interview is a fast, cheap and well validated way to ask a very large number of people something, and it beats waiting six weeks for fieldwork.
If they could not, you have to watch. What they will wear next spring, which launch gets copied, whether a price rise reads as expensive or as finally credible. People are not hiding those answers. They do not have them, and the only record is what they go on to do in front of each other.
Most of what a fashion or consumer brand decides in a given week falls in the second group. That is the reason we built the company on observation, and the reason the agents exist: a prediction about behaviour is only useful if something acts on it before the sales data confirms it.
The offer
Bring the audience you know best
Name the audience you understand better than anyone, the one where you would spot a wrong answer immediately. We will show you who moves it, what they are picking up and what they are asking for, with the real accounts behind every claim. Where it disagrees with you, one of us is wrong, and you will be able to check which.
Related reading: Hugo vs Aaru · AI agents for consumer brands · How we map out any audience · How accurate is AI consumer research · Hugo vs traditional market research · All comparisons
Frequently asked questions
What does Simile do?
Simile builds AI simulations of real people, which it calls agentic twins, and lets organisations survey them at scale. It was founded by Joon Sung Park, the lead author of Stanford's 2023 generative agents paper. The method comes from a 2024 study in which 1,052 Americans were each interviewed for two hours and a language model agent was built from every transcript. Those agents matched their human's General Social Survey answers at roughly 85 percent of the accuracy the humans themselves reached when retaking the survey two weeks later. Simile has reportedly raised 100 million dollars led by Index Ventures and 200 million dollars at a two billion dollar valuation led by Greenoaks, and names CVS Health, Gallup, Deloitte, Wealthfront, Itaú and Suntory as customers.
How accurate is Simile, and what does the 85 percent figure measure?
The 85 percent figure measures how closely a twin's survey answers match the survey answers of the person it was built from, divided by how closely that person matches their own answers two weeks later. It is an honest way to report agreement, because people are not perfectly consistent with themselves. What it measures is agreement between two self-reports. It does not measure purchases or any other observed behaviour, so it tells you how well the twin has learned what a person says about themselves.
Is Hugo a Simile alternative?
Hugo is the better choice when a consumer brand needs to predict behaviour rather than opinion: what to make next, how a launch will spread, which creator carries a product, what a price change does to different communities. Hugo observes real people in public, fits a behavioural model to each of them, and simulates influence across the follow graph that connects them, with no language model in the behavioural layer. Simile is the better choice for attitudes and stated choices, for private behaviour such as health or finance, and for any question that needs a recruited, representative panel.
What is the difference between a digital twin built from an interview and an observed behavioural model?
A twin built from an interview is a language model grounded in what a person said about themselves, so it reproduces how that person describes their own choices. An observed behavioural model is fitted to what a person actually posted, wore and reacted to, and outputs a predicted reaction directly. The distinction matters because decades of research show people cannot accurately report the causes of their own choices, and stated intentions explain only around 28 percent of the variance in later behaviour. In fashion, where imitation and status drive most decisions, the part people cannot report is the part a brand needs.
How do Simile, Aaru and Hugo differ on consumer simulation?
They build the population three different ways. Simile interviews recruited people and builds a language model twin of each one. Aaru reconstructs a population statistically from records such as census data, transactions, visits and search demand. Hugo observes real people in public, reached by chain referral through their actual connections, and models each person's behaviour and the influence between them on a graph. Simile and Aaru deliver predictions from a simulation you run. Hugo builds agents on top of its simulation and prediction infrastructure that work on a brand's decisions continuously and report into its Slack.
Does Hugo use a language model to simulate consumers?
No. Each person in a Hugo audience carries a small behavioural model fitted to what that person actually did, and influence between people is fitted from spread that was observed on the real follow graph. A language model is used only at the end, to phrase a finding for a human reader. In Hugo's published simulation work, a language model given the same information as the behavioural model came second at predicting who adopts a trend.
What are the agents Hugo builds on top of its simulations?
An agent is Hugo's prediction and simulation infrastructure pointed at one commercial decision for one brand. Agents are custom built and deployed for each brand by Hugo's engineers, who work inside the brand's team. The ones brands ask for first are the catalogue agent (what to push from what you already stock), the product agent (what to make next), the creator agent, the competitor agent, the quality agent, and the simulation agent, which tests a concept, price or campaign on a model of the real audience before launch. They run continuously and report into the brand's own Slack.
What can Simile do that Hugo cannot?
Hugo only sees people who are visible in public. It cannot reach private behaviour such as medication, banking or workplace attitudes, it cannot provide a sample with a documented recruitment frame for a regulator or a board, it under-represents people who do not post, and it does not ingest a customer's first-party data. A recruited panel with interview-based twins covers all four.