Files, email, web pages, and tool output land in your agent's context window carrying exactly the same authority as your system prompt. ContextDisarm takes that authority back — returning clean content, a signed provenance label, and an explicit envelope of what the agent is permitted to do with it.
Every operating system since the 1970s draws a hard line between code and data. Language models don't. A paragraph of retrieved text and the system prompt you spent a week tuning occupy the same address space and carry the same weight. Which means the moment your agent reads an untrusted file, whoever wrote that file becomes a co-author of your prompt.
A classifier hands back 0.71 and walks away. Your agent still has to decide what 0.71 means, and it will decide wrong at exactly the wrong moment.
Injection is natural language. It paraphrases, translates, encodes, and splits across turns. Pattern matching is permanently one rewrite behind.
A verdict that no component honors is a log line. Unless something in the execution path refuses the call, you have documentation of your own breach.
Traditional content disarm rebuilds a working file for a human to open — a hard problem across two hundred formats. Your agent never opens the file. That makes this a different, and smaller, problem: return the meaning, prove where it came from, and bound what may be done with it.
The document re-expressed as structured text and JSON, with active content, hidden layers, invisible characters, and embedded instructions stripped — plus an itemized list of exactly what was removed and why. Re-expressed, not reconstructed.
Every chunk leaves signed: where it came from, how far it has been transformed, which trust tier it belongs to. Labels survive chunking, embedding, and retrieval — so a document ingested six months ago is still marked when it lands in a prompt today.
A machine-readable statement of what this content is cleared for. Summarize, yes. Extract fields, yes. Become a tool argument, write to long-term memory, trigger an email — no, and here is the signal that closed the door.
from contextdisarm import Disarm disarm = Disarm(policy="finance-agent") ctx = disarm.file( "invoice-4417.pdf", source="email:vendor-inbox", ) ctx.text # clean, re-expressed content ctx.trust # "external.unverified" ctx.removed # [hidden_text, script, remote_image] ctx.allows("summarize") # True ctx.allows("send_email") # False
This is the part most of the category leaves to you. ContextDisarm ships middleware that sits between your agent and its tools and refuses any call whose arguments trace back to content that was never cleared for it. The envelope is not advice. It is a gate.
from contextdisarm import guard @guard(agent="finance-agent") def send_email(to: str, body: str): ... # …later, inside the agent loop, the model decides to act on the invoice send_email(to="ap@northwind-billing.co", body=ctx.text) ContextDisarmViolation: 'send_email' is not permitted for context ctx_01J8QK7M3NPX trust = external.unverified signals = [prompt_injection, hidden_text, exfil_beacon] allowed = summarize, extract_fields, answer_question # the call never leaves the process
Derived text inherits the trust tier of its source. A summary of an untrusted invoice is still untrusted, and the guard knows it three hops later.
Content that never passed through disarm has no envelope, so high-consequence tools refuse it by default. Silence is not consent.
Signed, immutable, exportable. When your auditor asks how the agent was prevented from acting on an untrusted document, you have the record.
LLM firewalls inspect the model call. Traditional CDR rebuilds files for humans. Classifiers return a number. ContextDisarm governs the lineage of the content itself — and stays useful alongside all three.
| Injection classifiers | AI gateways / LLM firewalls | File CDR | ContextDisarm | |
|---|---|---|---|---|
| What it inspects | Prompt and completion text | The model call, in transit | File structure and bytes | Content, and its lineage into the context window |
| What it returns | A probability | Allow / block on the request | A rebuilt file of the same type | Clean content + signed label + action envelope |
| Survives retrieval | No — evaluated per call | No — evaluated per call | n/a | Yes — labels persist through chunking and embedding |
| Enforced downstream | You implement it | At the model boundary only | n/a | At the tool boundary, by shipped middleware |
| Built for | Chat safety | Central governance and spend | Humans opening attachments | Agents that take actions |
Injection detection is commoditizing, and the only honest answer to that is measurement. We maintain an open corpus of injection and evasion payloads — hidden text, homoglyph and bidi tricks, encoding chains, multilingual rewrites, multi-turn setups, tool-output poisoning — and publish per-category results on every release, including the categories where we lose.
Measured by cd-bench run against a 75-case attack corpus and a 46-case clean
corpus of real-looking invoices, contracts, résumés, support threads and accessibility
patterns. Regenerated on every release; never typed by hand.
The corpus is ours. We wrote the attacks and the expectations, so a perfect
score is internal consistency before it is a claim about the world — a regression bar, not
a survey. The honest counterweight is mutation fuzzing, and we publish it: re-encoding a
caught payload drops detection to roughly 60%. That gap is the roadmap.
Asking a security team to ship untrusted enterprise documents to a third party in order to make them safe is a hard sell. So the disarm engine runs wherever you need it, and only telemetry and policy sync with us.
Wrap any MCP server in one command. Every tool response passes through disarm before it reaches the model — no changes to your agent code, no changes to the upstream server. Tool output is where injection actually lands; this closes it in an afternoon.
npx @contextdisarm/shield \ --wrap "npx @acme/support-tickets-mcp" \ --policy support-agent.yaml
Python and TypeScript. The engine runs inside your VPC; document content never crosses your boundary. Policies and detection models sync on your schedule.
One POST /v1/disarm away. Regional processing, configurable retention, nothing stored by default.
The invoices, contracts, resumes, claims, and support attachments your agents read every day were written by people outside your company. They arrive through channels you don't control, and they go straight into a system that treats text as instruction. That is the first place to put a checkpoint.
An agent that reads invoices and touches a payment system is one hidden paragraph away from rerouting a wire.
Attachments from anyone on the internet, parsed by an agent with access to account tooling and a send button.
Poison a document once and it sits in the index, retrieved on demand, for as long as the corpus lives. Labels are how you find it later.
We're onboarding a small number of teams running agents against untrusted documents and tool output. Bring a real pipeline and we'll disarm it with you.