Open data · CC BY 4.0

Buying-intent hard negatives

158 public posts that read as buying signals and are not. Every one was surfaced as a candidate by an automated scorer, then read and rejected by a person, with the reason recorded.

There are no positives in it. That is deliberate. A balanced corpus of obvious buyers and obvious noise is easy to assemble and teaches a classifier nothing it did not already know. What is scarce, and what breaks intent detection in practice, is the near miss: the post that carries every surface signal of a buyer and none of the intent.

What’s in it

  • ·Hashed post identifiers — SHA-256 of the original URL. No raw URLs, no titles, no body text.
  • ·Source platform — rss / reddit / hn (the subreddit / feed name is preserved as a coarse category).
  • ·Intent label — not_buyer on every row. The label is a person’s verdict, not the machine’s. A label produced by our own scorer would only teach a model to imitate us, and any accuracy measured against it would be accuracy measured against ourselves.
  • ·Rejection reason — why the reviewer said no. too_vague (65), not_relevant (61), wrong_product (32).
  • ·Panel agreement — whether our three judges were unanimous or split before a human ever saw the post. 141 unanimous, 17 split. Recorded separately from the label precisely so the two can be compared.
  • ·Age at capture and ISO week — coarse buckets, not timestamps. A timestamp plus a community narrows a post to one author.

The 32 that are worth your time

Thirty-two rows are marked wrong_product. Those authors were genuinely in the market — naming two vendors they were choosing between, stating a team size, saying the budget was approved. Real buying intent, for a different category.

They are the hardest case in intent detection and the one most systems get wrong, because every surface signal a keyword or a classifier looks for is present and correct. The only thing that disqualifies them is what they wanted to buy.

What is deliberately not in it: our scores. No component breakdown, no composite, no per-judge numbers. Those columns would let anyone recover the scoring formula by regression — four rows is enough once you take logs — and they are the one thing here that is ours rather than the internet’s. They are also of no use to a researcher: what makes a corpus citable is the labels and the observable properties of the posts, not one vendor’s arithmetic.

What’s NOT in it

  • · No raw post text, titles, or URLs.
  • · No author handles, follower counts, or any field that could re-identify a poster.
  • · No operator-level data (which SignalPipe customer scored which lead).
  • · No anchor sentences or product configurations.

Before you use it

  • This is not a base rate. Every row already cleared an automated gate before a human saw it. The share of posts in a raw feed that look like this is far smaller, and nothing here supports a claim about how common buying intent is in public communities.
  • One reviewer, one product’s perspective. The labels are one operator’s judgement about whether a post was worth answering for their product. A different product would relabel some of these — most obviously the 32 marked wrong_product, who were buyers for someone. Treat the label as “not a buyer of this”, never as “not a buyer”.
  • One labelling standard, one window.Ten consecutive weeks, 11 communities. Earlier data exists and is deliberately excluded: the bar for what counted as worth answering moved partway through the year, and blending the two would produce an average of two incompatible standards.

Get the dataset

signalpipe-hard-negatives-v1.csv — 158 rows, 21 KB, CC BY 4.0. No signup, no email. A permanent DOI on Zenodo will follow; the file above is the same data and is available now. Questions, corrections or a disagreement with one of the labels: contact@signalpipe.io.

Citation

SignalPipe (2026). SignalPipe Buying-Intent Corpus.
Released under CC BY 4.0. https://signalpipe.io/dataset
DOI: pending Zenodo upload

License

Released under Creative Commons Attribution 4.0 International (CC BY 4.0). You can copy, redistribute, remix, and build upon the material for any purpose, including commercial — as long as you give appropriate credit and indicate if changes were made.

See methodology for how the scores are computed, or /signals for the weekly aggregate trend digest.