Blog / Market research

What social listening misses: the 98% nobody measures

Social listening finds the sentence where someone named you. We went and measured how much consumer behaviour never gets written down at all. In our data it was 98% of it.

By the Hugo team · Updated August 17, 2026 · 6 min read

The number that started this

We spent most of this year building a map of what a real consumer audience wears. Not a trend report, a measurement: 876 New York accounts, 4,464 posts, every visible garment detected and cut out of frame.

One number from that run reframed how we think about the whole category. 2% of the frames were outfit posts. Two percent. Everything else we could see people wearing appeared incidentally: a jacket in a video about someone's apartment, a friend in the background of a restaurant clip, a bag on a stranger walking past.

Social listening, as the tools are built today, works on the 2%. It searches text for a keyword, finds the posts where somebody wrote about the topic, and analyses those. Everything in the other 98% is not ranked low or scored badly. It is not in the index at all.

Three things the method cannot see, by construction

It only sees behaviour that got written down

A customer who wears your jacket four times a week and never posts about it does not exist to a listening tool. Neither does the one who considered you and bought the competitor. The written mention is a rare event that sits on top of the behaviour, and treating the mention as a proxy for the behaviour assumes a stable ratio between them. There is no reason to think that ratio is stable across products, categories or demographics.

It over-samples whoever performs loudest

Keyword and hashtag retrieval returns the accounts that post the most about a topic in the most findable way. Those people are, by definition, atypical. If you optimise a range against them you have optimised against the people who talk about fashion rather than the people who buy it. We built our sampling around this specifically, reaching people through their real connections instead of through search, because finding an audience by keyword returns the performers, not the population.

It finds you only when someone names you

Retrieval is by brand name and keyword. The conversations where a customer describes exactly your category, exactly your problem, and never types your name are the most commercially valuable text on the internet, and they are invisible to a query that starts from your name. That is where the real objections and the real comparisons sit.

What social listening is genuinely good at

This is not an argument that the category is worthless. It is an argument about which question you are asking.

If the written mention is the thing you want, listening tools are the right tool and there is no better one. Reputation monitoring. Crisis detection, where you need to know in an hour. Tracking how a campaign landed. Reading complaint language verbatim so your copy uses the customer's words instead of yours. All of those are questions about what people said, and social listening answers them well.

It stops being the right instrument the moment the question becomes what will this audience want next. That answer is not in the text, because the behaviour that predicts it was never written down.

The alternative is observation, not better search

The fix is not a smarter query. It is inverting the order of operations.

Listening starts with a keyword and finds whoever used it. The alternative starts with a population and measures everything it does. You build a representative sample of a real audience first, then read every post those people make, whether or not the post is about your category, and extract behaviour from it directly.

In our case that meant computer vision on the frames rather than search over the captions, which is how you get to read the 98%. The output is different in kind. Not "3,200 mentions, 62% positive sentiment", but 832 distinct garments, black at 28.8%, white at 19.2%, with black taking over half of jackets while shoes flip to white. One of those is a summary of what got said. The other is a brief you can plan a range against.

What this costs you

Honesty about the trade, because observation has its own problems and we would rather state them than get asked.

It is much harder to build. Search is cheap and observation needs a sampling method, a vision model and an accounting of your own error. Ours is imperfect and published: our detector scores F1 0.597 on our own benchmark, the shipped graph is 37.6% usable nodes, and a filter we assumed was harmless was quietly removing 18% of garments, unevenly, until we caught it by re-running with it switched off.

It is also narrower per run. A listening tool covers every platform and every keyword at once. An observed map covers one audience properly. That is the trade we chose, because a shallow read of everyone has not, in our experience, ever settled a product decision.

The question to ask about any consumer data

One question separates most tools in this space: what would have to happen for a customer to show up in your data?

If the answer is "they have to write something, and it has to contain a word I searched for," you are looking at a small and unrepresentative slice, however large the mention count. If the answer is "they have to be in the audience, and be visible," you are looking at something closer to the population.

We publish the full method, including the parts that do not work yet, in mapping what consumers actually wear. The short version of why we build it this way is on what Hugo does for consumer brands.

Frequently asked questions

What are the limitations of social listening?

Three structural ones. It only sees behaviour that got written down, so silent use of a product is invisible. It only sees people who post, which over-samples the loudest rather than the typical customer. And it retrieves by keyword, so it finds you only when someone names you, missing every conversation where a customer describes your category without using your brand name.

What percentage of consumer behaviour does social listening capture?

A small fraction. In our own garment mapping work, only 2% of the frames we processed were posts about clothing. The remaining 98% of the clothing we could observe appeared incidentally, in the background of posts about something else entirely. Text-and-keyword social listening would have seen the 2% and missed the rest.

What is the alternative to social listening?

Observation rather than retrieval. Instead of searching for mentions of your brand, build a representative sample of a real audience, then measure what they do and use across everything they post, whether or not the post is about your category. That turns the question from what are people saying about us into what are these people actually doing, which is the question most product decisions need answered.

Is social listening still worth doing?

Yes, for what it is good at: reputation monitoring, crisis detection, tracking a campaign's reception, and reading complaint language in customers' own words. Those are all cases where the written mention is the thing you want. It stops being the right tool when the question is what an audience will want next, because that answer is not in the text.

Measure the 98%, not the mentions.

Bring an audience. We build the cohort, read what they actually do, and show you the evidence behind every number, including the ones we would rather not show you.