What we have learned

Findings from running a buying-intent scoring engine across Reddit, Hacker News and RSS feeds. These are the things that surprised us, written so they can be quoted without being wrong.

Qualitative on purpose. We publish a figure only where we can point at the run that produced it, and no individual post, author or URL is quoted here. Free to cite with attribution.

A buyer who never names the product, read correctly in six languages

The finding

Six posts were written in English, Spanish, Portuguese, French, German and Finnish, then judged three separate times each: 108 judgements against labels fixed before the run started. Two of the three buyer posts were identified in all 36 of their judgements. One of those two contains no product word anywhere — no "editor", no "software", no "tool" — only somebody describing the three hours an evening they lose cutting silences out of their own videos. No keyword filter can see that post, in any language.

  • ·The posts were written for the test rather than collected from anywhere, which is the point of doing it this way. The labels are certain because we set them, and every language contains the same six posts, so a difference in the result is a difference in reading rather than a difference in what happened to be posted that week.
  • ·The failures are the more useful half. Every false positive in the run — nine of them — was the same post: somebody celebrating that a video they had edited themselves passed a hundred thousand views. It was mistaken for a buyer in two of three English runs and three of three German ones. The listicle and the post from a person building a competing product were never once mistaken for buyers.
  • ·That failure is worth arguing with rather than defending. The person celebrating genuinely does edit video, so they sit squarely inside the audience a vendor wants to reach; they are simply not buying anything today. Separating those two states is the entire problem, and it is where this system is weakest.
  • ·One buyer post was read correctly in English, Spanish and Portuguese and missed in French, German and Finnish — identically in all three runs. The consistency rules out chance, but not translation. We wrote those translations ourselves, and a flat result in a language nobody here reads well cannot fairly be blamed on the model. It is recorded as unresolved rather than reported as a language weakness.

We screened six communities by hand. None of them contained a buyer.

The finding

Across six public communities and 72 posts read individually by the panel, zero were people trying to buy the thing being sold to them. The posts were on-topic, often highly engaged, and written by exactly the demographic a vendor would target. They were still not buyers.

  • ·The communities were chosen as the most plausible places to find that buyer, not as strawmen. Two were the highest-volume feeds already in production.
  • ·The three posts that initially scored as buyers were read again and rejected: one was somebody building a competing product, one was about error monitoring in a different domain, and one was looking for beta testers.
  • ·This is the clearest form of a pattern we keep meeting: the communities that discuss a category hardest are the ones whose members build in it rather than buy in it.
  • ·The practical consequence for anyone doing this: where you listen decides your results far more than how you score, and it is much cheaper to test a community than to tune a model.

Independent readers disagree about most posts, and that is normal

The finding

Across 255 scored posts, three judges reading the same post separately produced a mean spread of 0.478 on a 0-1 scale. Treating any disagreement as noteworthy flagged 76.5% of posts; only the top decile, above 0.75, is unusual enough to be worth a human glance.

  • ·Mild disagreement is the ordinary state of independent assessment, not a warning sign. A system that reports every disagreement is telling you nothing, in the same way a system that reports none is.
  • ·We had the threshold at 0.30 and it fired on three calls in four. It is now at 0.75, chosen from this distribution rather than from intuition.
  • ·The general lesson is about confidence signals of any kind: a flag that fires on the majority of cases trains the person reading it to ignore it.

Twelve of twenty-seven feeds produced nothing at all

The finding

Of 27 active sources running against one product, 12 returned no usable post in 14 days and several had returned none ever. Removing them cost no leads and returned roughly 44% of the fetch budget to the feeds that were working.

  • ·Three separate rows were pointed at the same URL, so one feed was being fetched three times a cycle while sharing the penalty for its own noise across three counters.
  • ·Dead sources are not free. Anything polling on a schedule spends its budget on empty feeds exactly as readily as on productive ones.
  • ·After the removal the share of the queue coming from the single noisiest community fell from 52.9% to 10%, and total volume was unchanged.

On-topic is not in-market

The finding

The communities that discuss a product category most intensely are usually among the worst places to find that category's buyers, because the people discussing it are largely the people building in it - and builders tend to build their own rather than buy.

  • ·This is the most common reason keyword-based lead tools disappoint: volume of relevant-sounding discussion is a poor proxy for demand.
  • ·It holds across both products we run the engine for, and it is a property of the audience rather than of the scoring.
  • ·The practical consequence is that choosing where to listen matters more than tuning how you score.
  • ·Communities of operators - people running a business function day to day - behave differently from communities of makers, and are the ones worth listening to.

Popularity is anti-correlated with buying intent

The finding

A post's upvotes and comment count measure reach, not whether its author wants to buy anything, and in business communities the two often point in opposite directions: the most-upvoted post is usually a success story, and a success story is not a buyer.

  • ·Engagement is an appealing ranking signal because it is cheap and always available. It is also close to irrelevant to intent.
  • ·We removed engagement from the path that decides which voice a reply is written in after measuring this.
  • ·Any lead tool that sorts its queue by popularity is sorting away from its buyers.

The best signals are unglamorous

The finding

A short, low-engagement post stating a specific requirement - an existing operation, a named constraint, a budget - outperforms a longer, more articulate, more popular post on the same topic almost every time.

  • ·Specificity about the author's own situation separates buyers from commentators more reliably than enthusiasm about the category does.
  • ·"Where do I start" is not a buying signal, however genuine: no stated requirement, no existing operation, nothing to sell to yet.
  • ·This is why a strict filter that returns a short queue is doing its job rather than failing at it.

Telling a buyer from a bystander is not a keyword problem

The finding

Separating "someone discussing this problem" from "someone trying to solve it right now, for themselves" is not achievable with keyword matching or a similarity threshold, because both populations use nearly identical language.

  • ·Substring matching in particular fails in both directions - matching inside unrelated words, and missing intent phrased in the author's own terms.
  • ·A single language-model pass is unreliable exactly where it matters most: the ambiguous middle, where a confident answer is worth least.
  • ·Independent assessors allowed to disagree, with the disagreement surfaced rather than averaged away, is the approach we settled on.

For how the scoring works, see methodology.

Want the underlying anonymized dataset? /dataset.