There are only two kinds of AI in a fashion business
The first kind generates something from a brief you already wrote. Product imagery, on-model shots, marketing copy, a first-pass print. You know what you want, the model makes it faster and cheaper than a photoshoot.
The second kind decides something. What to make, in which colorway, in what quantity, for whom. Here the model is not producing an asset, it is producing a claim about the world, and the claim is only as good as what the model was allowed to look at.
Nearly every honest evaluation of AI in fashion collapses to that split. The first group is largely solved and getting cheaper every quarter. The second group is where the money is and where most tools are much weaker than the demo suggests.
Group one: generation. Mostly solved.
1. Product and on-model imagery
Turning a flat-lay into a professional model shot now takes seconds instead of a shoot day. H&M has publicly used AI twin models to scale product visuals across markets. This is the most mature use case in fashion AI and the one with the clearest cost line: the work was expensive, it is now cheap, and the quality gap closed faster than most people expected.
How to evaluate it: on cost per asset and how well it holds your brand's look across a full range. The underlying models are close enough that the differentiator is workflow, not intelligence.
2. Copy, product descriptions and localisation
Thousands of SKUs, each needing a description in six languages. This is a volume problem, and volume problems are what these models are best at. Commoditised.
3. Virtual try-on and sizing
Nike Fit scans a customer's feet through a phone camera to size a shoe. Try-on and fit prediction reduce the single largest cost in fashion ecommerce, which is returns. Genuinely valuable, genuinely working, and it belongs in this group because the model is transforming an input you already have, the customer's own body, rather than making a claim about a market.
Group two: decision. Only as good as the data underneath.
4. Trend detection
Shein's proprietary system detects emerging demand in real time and triggers small-batch production against it. That loop, detect and produce in weeks, is the actual competitive advantage, and the detection half is the harder half.
The critical question for any trend tool is the population it reads. A model trained on runway shows, fashion editorial and hashtag search will tell you what the fashion industry is currently discussing. That is real information, and it is a lagging indicator of the industry, not a leading indicator of demand. A model that reads what ordinary people are wearing is answering a different question.
We measured this gap directly. In our own mapping work, only 2% of the frames we processed were outfit posts. The other 98% of the clothing we could see was incidental: someone wearing a jacket in a video about their apartment, a friend in the background of a restaurant clip. Any system that reads fashion content is reading the 2%, and the 2% is the part that is performed.
5. Demand forecasting and allocation
Burberry uses AI in supply chain management to redistribute inventory against real-time demand signals, cutting markdowns and excess stock. This works because the signal is first-party: their own sell-through, their own stores, a closed loop where the model can be scored against what happened.
It works much less well for the thing brands actually want, which is knowing demand for a product that does not exist yet. Your sales data cannot tell you about the customer who looked and left, or the colorway you never made. That is the limit we wrote about in what your bestsellers can't tell you.
6. Assortment, colorway and range planning
The most valuable and least served use case. Deciding that a jacket should be black rather than navy is a decision made thousands of times a season, usually on taste and last year's numbers.
It is answerable with measurement. When we read 832 distinct garments off 876 real New York accounts, black came out at 28.8% of everything worn and white at 19.2%, with 58.3% of the wardrobe carrying no chromatic color at all. More usefully, the aggregate hides the split that matters: black takes over half of jackets and eyewear, while shirts and shoes flip to white, and denim owns exactly one category, pants. An aggregate palette is a fact. A palette per garment type is a plan.
7. Competitive and share-of-wardrobe tracking
Most competitive tooling reads competitors' published output: their site, their drops, their prices, their marketing. Useful, and it tells you what a competitor said. It does not tell you whether anyone is wearing it. Share of a real wardrobe, measured on bodies rather than inferred from earnings calls, is a different and harder measurement, and it is where the comparison customers are actually making shows up first.
The question to ask a vendor
For group one, ask about cost per asset and brand consistency. The models are close to interchangeable.
For group two, ignore the model entirely and ask three things about the data:
What population does it observe, and how was that population reached? Hashtag and keyword search returns whoever performs a topic loudest, which is a biased sample in a specific and uncorrectable direction. Sampling through real social connections returns something closer to the actual population, with a bias you can at least name.
Can it see garments on ordinary people, or only in fashion content? This is the 2% question above, and it is the single biggest difference between two tools that describe themselves identically.
Does the vendor publish its own error rate? Anyone can produce a confident chart. We publish ours: our garment detector scores F1 0.597 on our own benchmark, our shipped graph is 37.6% usable nodes, and a filter we thought was harmless was quietly deleting 18% of garments unevenly by category until we caught it. Those numbers are unflattering and they are the reason the outputs can be trusted. A vendor who will not tell you where their pipeline fails is not a vendor who has measured it.
Where this is going
The generation half of fashion AI is finished as a competitive question. Everyone will have it, it will cost almost nothing, and no brand will win on having better AI product photos.
The decision half is barely started, because it depends on a substrate almost nobody has built: a structured, measured, error-characterised picture of what real people wear, that a model can query. That is the work we publish in mapping what consumers actually wear, and it is the reason we describe Hugo as infrastructure for fashion brands rather than as another trend tool.