A buyer who never names the product, read correctly in six languages
Six posts were written in English, Spanish, Portuguese, French, German and Finnish, then judged three separate times each: 108 judgements against labels fixed before the run started. Two of the three buyer posts were identified in all 36 of their judgements. One of those two contains no product word anywhere — no "editor", no "software", no "tool" — only somebody describing the three hours an evening they lose cutting silences out of their own videos. No keyword filter can see that post, in any language.
- ·The posts were written for the test rather than collected from anywhere, which is the point of doing it this way. The labels are certain because we set them, and every language contains the same six posts, so a difference in the result is a difference in reading rather than a difference in what happened to be posted that week.
- ·The failures are the more useful half. Every false positive in the run — nine of them — was the same post: somebody celebrating that a video they had edited themselves passed a hundred thousand views. It was mistaken for a buyer in two of three English runs and three of three German ones. The listicle and the post from a person building a competing product were never once mistaken for buyers.
- ·That failure is worth arguing with rather than defending. The person celebrating genuinely does edit video, so they sit squarely inside the audience a vendor wants to reach; they are simply not buying anything today. Separating those two states is the entire problem, and it is where this system is weakest.
- ·One buyer post was read correctly in English, Spanish and Portuguese and missed in French, German and Finnish — identically in all three runs. The consistency rules out chance, but not translation. We wrote those translations ourselves, and a flat result in a language nobody here reads well cannot fairly be blamed on the model. It is recorded as unresolved rather than reported as a language weakness.