Spaces:
Sleeping
Robustness scorecard: normalization closes homoglyph and zero-width evasion, but not paraphrase
Clean-data accuracy is not deployment accuracy. This demo scores a message for fraud, lets you run an attack a real fraudster would use, then applies a defense and shows whether the detector recovers.
I ran it as a scorecard over 819 fraud lures from LureBench's core test set. Evasion rate = of the lures a detector caught on clean text, the fraction that slip below threshold after the attack, measured raw and after input normalization.
The pattern that matters:
- Homoglyph and zero-width: normalization drives evasion to 0% for both detectors. These attacks are losslessly reversible.
- Leet: mostly recovered, with a 16% residue where the digit 1 is ambiguous between i and l.
- Whitespace and semantic paraphrase: untouched. Re-joining split words or undoing a paraphrase would corrupt real text.
So the typographic robustness gap is easy to close and the semantic one is the real work. The full API also exposes the content-safety models teams actually deploy (Llama Guard, OpenAI moderation, an LLM judge); the underlying benchmark separately shows Llama Guard at 0% true-positive on romance-baiting lures while catching other scams.
Full writeup: https://dev.to/immu4989/a-fraud-classifier-at-96-recall-and-the-one-character-edit-that-walks-the-lure-through-36kb
Code and the reproducible scorecard: https://github.com/immu4989/lurescope
Try the demo above and tell me where it breaks, especially on the semantic attacks.
