1. The problem keyword alerts can’t solve
Suppose you sell a tool that automates Reddit lead discovery. A keyword alert for “Reddit leads” will fire on every post that mentions Reddit and leads, regardless of whether the author wants to buy something or sell something. The signal-to-noise ratio collapses fast.
A lexical match doesn’t know the difference between “I’m looking for a tool to find Reddit leads” (a buyer) and “5 ways our tool finds Reddit leads” (a competitor’s blog post). A semantic match does. That’s the entire reason SignalPipe exists.
2. Anchor sentences as semantic targets
Each product configures 5–10 anchor sentences: short, natural-language examples of the buying intent it wants to detect. They’re written in the buyer’s voice, not the seller’s. Examples:
- “I need a tool to monitor Reddit for sales leads”
- “Looking for an AI sales agent that can find prospects automatically”
- “My outbound process is too manual and I want to automate lead discovery”
Your anchors are embedded once and cached; every incoming post is embedded the same way and matched against the closest one. How well it matches becomes the starting point of the score — not the whole of it.
Why anchors instead of a single product description? Because real buyers don’t describe your features — they describe their pain. Several short anchors covering different framings of the same intent outperform a single long product blurb on every dataset we’ve tested.
3. Semantic match is necessary, not sufficient
A post can match an anchor almost perfectly and still be worthless to you: a stale repost from eight months ago, a sarcastic aside, a rhetorical question in someone’s blog promotion. Meaning alone does not tell you whether replying is a good use of your morning.
So several independent things about a post are assessed and combined — not only what it means, but how urgent and how specific it is. Posts past an age cutoff are skipped outright. How recently it was written, whether anyone is paying attention to it and the author’s history only decide the order leads are shown in, never whether they get in, and whether the person writing it is a plausible buyer rather than someone marketing at the same audience you are is left to the judges.
The combination is deliberately unforgiving. A post that scores brilliantly on one dimension and badly on the rest does not get to average its way to the top of your queue — weakness anywhere pulls the whole score down. That is a design choice, and it is the reason the queue stays short.
The output is a 0–100 content score — what the post itself says, before anything about its context is taken into account.
4. The competitor-floor heuristic and its honesty cost
If a post mentions one of your configured competitor names and its author is complaining about it, comparing it or replacing it, the scout raises the score to a floor. A passing mention gets no floor, and the panel still judges every one. Why: someone publicly evaluating your competitor is one of the highest-intent signals there is, and you don’t want a weak embedding match to bury it.
The cost is real and worth stating plainly: a rule that surfaces competitor mentions will sometimes surface the competitor’s own marketing team talking about their own product. We do not pretend otherwise. Both numbers are kept — what the post itself said, and what it scored once context was applied — so you can always see when a lead is strong because of its content or because of its context, and we flag the ones where those two disagree sharply.
The draft is calibrated from what the post actually said, not from the context-adjusted number. So a misread competitor mention still reaches your queue, where you can judge it — but the reply stays helpful and value-first rather than being upgraded into a hard pitch aimed at somebody who was never going to buy.
5. It learns which of your sources are worth listening to
Every feed you listen on — a subreddit, a Hacker News search, an RSS URL — builds its own track record with you. A source that keeps producing leads you approve earns more of your attention over time. One that keeps producing noise earns less. The learning is per-feed rather than global, so a single bad source gets demoted on its own instead of quietly suppressing everything else you listen on.
Your rejections carry more information than your approvals, which is why the reason matters. A lead that was outright spam is a different failure from a well-targeted lead that happened to come from someone who is already your customer — the first says the source is bad, the second says the source is fine and the timing was off. Those are treated differently, so choosing the reject reason accurately is worth the extra two seconds.
The loop is deliberately asymmetric in one direction: missing a real prospect costs you more than skimming one extra mission you did not need. The system is tuned around that trade, not around keeping your queue tidy.
6. Three judges who disagree
Drafting runs through a panel of three judges with genuinely different dispositions — a skeptic, an analyst, and an optimist. They read the same candidate independently and they frequently disagree. That disagreement is the mechanism, not a bug to be smoothed over: a lead all three accept is worth your time, and one only the optimist likes usually is not.
The draft is then calibrated to how strong the signal actually turned out to be, judged from what the person wrote rather than from the final score. Someone openly shopping gets a direct answer. Someone asking a general question gets a genuinely useful reply that earns the mention instead of forcing it. The difference in practice:
- Closer — propose a specific concrete next step (demo, trial, link)
- Advisor — consultative, acknowledge the problem, introduce the product as a natural fit
- Educator — answer the question first, pitch only if it fits naturally
Char budgets are passed in too: 280 for Twitter replies, 500 for Reddit DMs, 300 for manual outreach. The swarm targets the budget at draft time — operators don’t inherit a draft that needs to be cut in half.
7. Failure modes we know about
Honest list. Most are mitigated; one or two are open.
- Sarcasm — embedding similarity can’t detect inverted intent. Mitigated by a sarcasm-detection pass on the brain before drafting.
- Stale reposts — same URL crossposted to 5 subreddits. Mitigated by the interactions ledger (unique on URL + product).
- Competitor-marketer false positives — discussed in §4. Surfaces in queue but doesn’t produce a hard pitch.
- Bot accounts and karma farms — open. Author-reputation component helps but doesn’t fully solve. Operator review is the backstop.
- Cross-product anchor leakage — a generic anchor like “I need a sales tool” will match too broadly. Best practice: write anchors as specific as the product warrants.
Frequently asked questions
Why use anchor sentences instead of keyword alerts?
Keyword alerts can’t distinguish a buyer from a competitor: a search for “Reddit leads” fires on both “I need a tool to find Reddit leads” and “5 ways our tool finds Reddit leads.” Anchor sentences are embedded once, and incoming posts are matched against the closest one — so the system reacts to the intent expressed, not the surface tokens.
Why is semantic matching alone not enough?
A post can match your anchors almost perfectly and still waste your time: a stale repost, a sarcastic aside, a rhetorical question inside somebody’s blog promotion. So meaning is one of several things assessed, alongside how urgent and how specific the post is. They combine unforgivingly — strength on one does not rescue weakness on the others. Popularity and recency only change the order leads are shown in, and whether the author is a plausible buyer rather than someone selling to the same audience is left to the judges.
What does the competitor floor do?
It gives a post about a competitor you configured a higher floor when its author is complaining, comparing or switching, because someone publicly evaluating your competitor is among the highest-intent signals there is. Both numbers are kept — what the post itself said, and what it scored once context was applied — so you can see when a lead is strong only because of context. The draft is calibrated from the post itself, so a misread mention still appears for you to judge but stays helpful rather than becoming a hard pitch.
How does it learn from my decisions?
Every feed you listen on builds its own track record with you. A source that keeps producing leads you approve earns more of your attention; one that keeps producing noise earns less. The learning is per-feed rather than global, so one bad source is demoted on its own instead of quietly suppressing your good ones. Your reject reason matters too, because an outright spam lead and a well-targeted lead from someone who is already your customer are different failures and are treated as such.
How does the judge panel decide what reaches me?
Three judges with genuinely different dispositions read each candidate independently and often disagree. That disagreement is the mechanism, not a flaw to be smoothed over: a lead all three accept is worth your morning, and one only the most optimistic of them likes usually is not — so it is dropped rather than forwarded. The draft is then calibrated to how strong the signal actually turned out to be, and written to fit the channel it will go out on, so you never inherit a draft that needs cutting.
What failure modes does SignalPipe acknowledge?
Five known: sarcasm (embeddings can’t detect inverted intent — mitigated by a sarcasm-detection unit-test pass), stale reposts (mitigated by the interactions ledger, unique on URL+product), competitor-marketer false positives (surface in queue but don’t produce hard pitches), bot accounts and karma farms (open — author-reputation helps but operator review is the backstop), and cross-product anchor leakage (mitigated by writing anchors as specific as the product warrants).