Most of the security industry has filed prompt injection under the same heading as spam and malware: a content problem. Feed the content to a classifier, get a verdict, tune the threshold. It is a comfortable answer, and it is the wrong answer.

The admission inside a detection rate

A filter answers one question: is this content malicious? If it is good, it answers correctly almost always. But “almost always” is the whole problem. Every detection rate is an admission of a miss rate, and against an attacker who can re-encode the same instruction in base64, in an obfuscated character set, across two sentences, or in a language the filter has never seen, the misses are where the value leaks.

That is not a criticism of detectors. It is a statement about what the problem is. An agent does not exfiltrate because some text looks malicious. It exfiltrates because it was told to, by text it treated as instruction, and it had the authority to do what it was told.

The different question

So the useful question is not about the content at all:

What authority does this action require, and does the session still hold it?

Authority is a property of the session, not of the bytes. It can be computed. It is a membership test, not a probability: does this host appear in the origin set derived from the repository’s own committed history? Does this package match what the manifest already names? Does this tool definition still match the one a person approved?

A membership test has no miss rate, because it is not guessing about intent. It is checking whether a finite, human-authored set contains a value. Either it does, or it does not.

Why “monotone” is the word that matters

The property that makes this sound rather than merely clever is that authority only ever goes down.

When untrusted content enters the context, whether a PR title, a fetched page, a tool result or an MCP description, the session’s authority is attenuated at the moment it arrives, not when it is interpreted. Encoding, obfuscation, paraphrasing: none of it matters, because the attenuation already happened before the model read a word.

And once attenuated, nothing inside the context can raise it back. An injected instruction cannot widen what the agent is allowed to do. The best an attacker can achieve is that the session asks for permission more often, and asking is the one thing that routes the decision to a human instead of to the bytes.

This is not a new idea. The research community settled the direction in 2025 with CaMeL’s capability-based information-flow control, and could not deploy it because it required owning the interpreter and the harness. The harnesses have since shipped the interception point, the hook that fires before every tool call. That is where the decision belongs.

What this means in one line

A scanner says “I will try to catch the bad content.” An authority decision says “this action is outside what you were granted, so it does not run.” The first is a rate. The second is a guarantee, and the difference is whether the leak already happened when the alarm goes off.


This is a product-direction perspective, not a claim that every connector or enforcement point is already implemented. Coverage must be shown for the environment that is actually connected.