\n\n\n\n Invisible Ink for the Unicode Age - AgntHQ \n

Invisible Ink for the Unicode Age

📖 4 min read•764 words•Updated Sep 7, 2026

You open your inbox on a Tuesday morning. There’s an email that looks like it came from your bank. The subject line reads clean. The body text reads clean. Your spam filter waved it through without a flicker, because as far as the filter could tell, there was nothing suspicious in it. And your filter was right, in the narrow sense that filters are ever right. It read every character it could see. The problem is that it couldn’t see all of them.

That’s ASCII smuggling, and according to Microsoft, spammers started leaning on it hard beginning in February 2026.

An attack that changed careers

What makes this one interesting to me isn’t the technique itself. Hiding characters inside text is not new thinking. What’s interesting is the migration path. ASCII smuggling got its reputation Invisible Unicode goes in, the model reads it as part of the prompt, and now your helpful assistant is doing something you never asked it to do.

Now the same trick is being pointed at email filters. Microsoft’s own telemetry, which surfaced out of its Defender for Office 365 prompt injection protection research, shows the crossover. The team looking for AI attacks found spammers instead.

I keep coming back to that detail because it says something uncomfortable about how security research works right now. A defensive effort aimed squarely at AI-specific threats ended up documenting a traditional phishing evasion campaign. The techniques don’t respect the categories we invented for them.

Why this keeps happening

Every text-processing system makes assumptions about what text is. Filters assume the characters they can parse are the characters that matter. Language models assume the tokens in their context window are tokens someone meant to put there. Humans assume that what’s rendered on screen is what’s in the file.

Unicode contains blocks that violate all three assumptions at once. Characters that render as nothing. Characters that exist in the byte stream, get processed by machines, and never reach your eyes. If your detection logic operates on the visible string, you are working from an incomplete copy of the document.

So the same gap gets exploited twice, against two entirely different targets, using the same mechanism. That’s not two vulnerabilities. That’s one vulnerability with two customer bases.

What this means if you’re buying AI tools

This is the part I care about, because it’s where the reviews I write run into reality. A lot of AI products are, functionally, a text pipeline with a model in the middle. Content moderation tools. Email assistants. Document summarizers. Agents that read your messages and take actions on your behalf.

Every one of those products inherits this problem, and most of the vendors selling them have not thought about it. I’ve sat through demos where the pitch was entirely about accuracy on clean input. Nobody demos adversarial Unicode. Nobody shows you what the agent does when the document contains instructions it wasn’t supposed to see.

If you’re evaluating tools right now, the question worth asking your vendor is not “how accurate is your model.” It’s “what does your input normalization look like.” Do they strip invisible characters before processing? Do they log what they stripped? Can they tell you what the raw bytes were versus what the model actually consumed?

Most will not have a good answer. Some will not understand the question. That’s useful information about how seriously they take the security side of what they built.

The uncomfortable read

The optimistic framing here is that Microsoft caught this, published it, and defenders now know to look. That’s real, and the visibility matters.

The less comfortable read is what the timeline implies. A technique gets attention Researchers study it in that context. Spammers, who have never cared about our taxonomies, notice it works on older targets too and start using it at scale. The gap between “known technique in one domain” and “widespread abuse in another” turned out to be short.

Which suggests the reverse trip is also available. Every evasion trick that already works against traditional filters is a candidate for pointing at models. Anyone building AI products on top of text they didn’t write should assume that catalog is being worked through right now, in both directions, by people who read the same research you do.

The invisible characters were always there. We just finally have two groups of attackers who found a use for them.

🕒 Published:

📊
Written by Jake Chen

AI technology analyst covering agent platforms since 2021. Tested 40+ agent frameworks. Regular contributor to AI industry publications.

Learn more →
Browse Topics: Advanced AI Agents | Advanced Techniques | AI Agent Basics | AI Agent Tools | AI Agent Tutorials
Scroll to Top