Blog / Method

We ran a study on 1,328 posts. 287 survived.

Every AI research tool will tell you how much data it processes. Almost none will tell you how much it throws away. Here is a real run through our pipeline with the funnel published, including the part where the answer was disappointing.

By the Hugo team · Updated August 5, 2026 · 7 min read

"We analysed 50,000 posts" is the most common claim in AI research marketing and one of the least meaningful. The number tells you what a scraper collected. It says nothing about how much of it was relevant, and relevance is the entire job.

So here is a run we did on our own infrastructure, with the numbers published rather than the headline one. The question we asked was deliberately hard: what do fashion buyers, merchandisers, demand planners and consumer insights professionals actually complain about in their work?

1,328 candidate posts gathered across six platforms. 287 kept after validation. 78% discarded before a single conclusion was drawn.

The four stages, and the one that matters

A grounded research system does four things. It turns your question into many specific search queries rather than one keyword. It runs those across platforms. It validates each returned post against the original question. Then it reasons only over what survived.

Stage three is the one that separates research from scraping, and it is the one most tools quietly skip. Skipping it is what lets a vendor claim a large number. It is also what produces confident findings built on job adverts.

What the discarded 1,041 actually were

This is the part worth being specific about, because it explains why the filter has to be so aggressive. A keyword match is not a relevance match. Searching professional topics reliably returns:

All of it matches the words. None of it answers the question. If those posts reach the reasoning stage, the model will faithfully summarise them and hand you a confident read on a market that is really a summary of recruitment advertising.

Platforms are not interchangeable

The same question returned wildly different yields depending on where it was asked. Validated posts by platform, from this single run:

PlatformValidated posts
Facebook91
TikTok83
LinkedIn55
Twitter/X24
Instagram23
Reddit11

Anyone selling you a single-platform tool is selling you a single-platform answer. The distribution also shifts by question: ask about a consumer category rather than a profession and the ordering changes completely.

The honest part: this run underdelivered

We could stop there and it would read as a clean methodology piece. It would also be misleading, because the interesting result was a negative one.

Of the 287 posts that passed validation, the genuinely useful ones, practitioners describing a real frustration in their own words, were a small minority. Most of what survived was still career content, industry commentary and people explaining the job rather than complaining about it. The pipeline worked. The population did not exist.

That is a real limit and it is worth naming plainly: professional B2B audiences mostly do not post their working frustrations in public. A buyer who is furious about a forecast subscription tells a colleague, not TikTok. Notice that Reddit, the platform where professionals are most candid, returned the fewest validated posts of any platform here, because Reddit's volume on this topic is simply small.

This is the same limit we would tell any customer about before they bought anything. Public consumer categories work well, because consumers narrate their preferences constantly and in public. Private professional populations do not, and for those, ten proper interviews beat any amount of social listening. We wrote about the same boundary in How accurate is AI consumer research? AI consumer research tools, compared Hugo vs traditional market research agencies.

What to ask any AI research vendor

If you are evaluating tools in this category, these four questions separate the grounded ones from the fluent ones quickly.

The underlying principle is simple enough to state in one line: if a finding cannot be traced back to something you can open and read yourself, it is not a finding. Applying that rule kills most impressive-looking AI research output. What survives is much smaller, and it is the only part worth making a decision on.

Why we published the disappointing number

Because the alternative is the thing we think is actually wrong with this category. The failure mode of AI research is not obvious errors, which you would catch. It is fluent, plausible answers that match your priors and are not connected to anything a real person said. That output is most convincing exactly when it is wrong, which is the worst possible property for something you are about to make a buying decision on.

The only defence is auditability, and auditability means showing the funnel, including the runs where the funnel produced less than you hoped. We benchmark our simulation output against real survey data for the same reason, and publish the results on our validation page.

Bring a question and watch the funnel run.

Pick a consumer question you actually need answered. We will run it live on the call, and you can see what gets thrown away as well as what comes back.