01 / AI & AGENT SECURITY
Know every agent.
Decide before it acts.
An inventory tells you which agents exist. It does not tell you what they reach through permissions they borrowed, and it cannot stop the next tool call. This module does both, and reports the empty case as empty.
THE LOOP THIS PRODUCT HAS TO CLOSE
EDGE + GITHUBingest
The only live inbound paths. Ten connector kinds are registered and fail closed.
BUILTgraph
Entities, edges and provenance, projected from events that really happened.
BUILT, THIN INPUTenrich
Classification and entitlement resolution over whatever the graph actually holds.
BUILTanalyze
Exposure scans, posture rules, path search, rollout simulation.
BUILT, NO BASELINEprofile
Per-entity reach and identity profiles. No behavioural or peer baseline.
ONE ACTIONremediate
Propose, approve, execute, verify. The only effector is revoking a session.
WHAT DOES NOT ARRIVE
Nothing discovers third-party AI applications on its own. The inventory is whatever the edge and GitHub put into the graph: agents, sessions, the authority decisions they produced, and the secrets and destinations a finding named. Identity, drive and SaaS connectors are registered against a fail-closed provider that answers NOT CONNECTED without a credential and UNAVAILABLE with one, so there is no usage analytics and no vendor-side discovery behind any of this. What is connected today ↗
WHAT CANNOT BE ACTED ON
The control plane has exactly one effector: revoking an agent session, which takes authority away and can never grant it. Marking a tool revoked changes our record of it. It does not reach a vendor, disable an account or remove anybody’s access, because there is no connector to act through. The enforcement that does exist is the hook, on the machine, on the next tool call.
A module is a drawing; a loop has to close. Zorro owns the middle of this one and is missing both ends, so read every capability below against the two cells above it. Discovery is exactly as populated as the graph is, and an unpopulated graph returns an empty inventory rather than a reassuring one.
[ 01 ]THE INVENTORY+
Four kinds of thing
that act on your behalf.
Discovery walks the graph for AI tools and agents, keeps whatever a person has already reviewed, and classifies only what is new. Nothing is overwritten by a rediscovery.
ENTITY KINDS THIS MODULE READS
ai_agentai_toolmcp_servermcp_toolaccountsessionsecretnetwork_destination
Every discovered entity keeps the provenance it arrived with: the source that observed it, the evidence_digest behind that observation, and the time it was recorded. An entry with no provenance is not an entry we are willing to show you.
[ 02 ]THE DECISION IN THE LOOP+
The inventory is the map.
The hook is the gate.
One binary reads a hook payload on standard input, answers on standard output and exits. No network, no daemon to reach, no lock to take, once per tool call, on the agent's critical path.
FIVE HARNESSESOne answer format each.
Claude Code, Codex, Cursor, GitHub Copilot and Gemini CLI each send their own payload and expect their own reply. Which one is calling comes from the flag, then the environment, then the shape of the payload itself.
SHADOW FIRSTThe default blocks nothing.
The first run on any machine records what it would have decided. That is how the origin sets get derived from what a repository actually does, and how the prompt rate is measured before anyone is interrupted. Enforcement is opt-in.
FAILURE ASKSNever a silent allow.
A payload that will not parse, a ledger that will not open, a tool or a provider that is not recognised: every one of them becomes a question for a person. The absence of an answer is not an answer.
ONE REFUSALA revoked session.
A person, or the control plane through the workstation daemon, stopped it. Nothing that session proposes proceeds, and the refusal lands on the next tool call rather than in a report.
THE MCP PROXY
caveat-mcp
on the stdio relay, before the harness sees the answer
Every tool listing is reduced to a pin per tool, a hash over the canonical JSON of the whole tool object plus the server's identity, recorded before the answer is passed on. The first observation is the baseline, so the hook can never see a definition the store has not. Tool descriptions, initialize instructions, results and sampling requests are all recorded as untrusted content, which means they can only narrow a live grant. The server also runs on an allowlisted environment rather than the harness's own: tokens in the parent are not the server's to read.
WHAT THE DECISION COSTS
caveat-bench, on recorded sessions
measured, with the caveat attached
111 to 640 µs p50, with p99 in milliseconds and machine-load dependent. The number we are not happy with is the other one: 36.66 mean and 25 p50 prompts per session on a frozen transcript set, against a target of two or fewer. That is the active engineering priority, and it is the honest reason shadow is the default rather than a setting we hope you change.
[ 03 ]CLASSIFICATION+
Sanctioned, shadow,
or honestly under review.
The posture the entity carries decides the state. Nothing infers approval from popularity, from usage volume or from a vendor's reputation.
FOUR STATES, RECORDED ON THE TOOL
under_review
The default for anything newly discovered. It is not a verdict, it is the absence of one, and it is recorded as such rather than defaulting to approved.
sanctioned
Approved by a named person. The review writes the reviewer, the time and the note onto the tool row, so the approval has an author.
shadow
Present and not approved. The exposure profile raises it as a finding: an unsanctioned tool connected without administrative review.
revoked
Approval withdrawn in the record. The row stays, because deleting it would delete the history of having approved it. Withdrawing approval is a decision, not an enforcement: nothing here reaches a vendor and turns the tool off.
The risk score attached to each state is an ordinal label for sorting a queue. It is set by the state, not derived from observed behaviour, and it is not a probability. Read the state; the number is only there to order the list.
[ 04 ]INHERITED REACH+
The permission it borrowed
is the one that matters.
An agent rarely holds access of its own. It runs as somebody, and it reaches whatever that somebody reaches, which is why an agent inventory that stops at the agent is not an answer.
WHO IT RUNS AS
Operating principals, read from edges.
Incoming uses and authenticates_as edges name the identities an agent acts through. It is a graph read, not an assumption about what a token probably carries.
WHAT IT REACHES
Traced, then counted by sensitivity.
The access path is walked to the files, data sources and databases at the end of it, and what is found is counted as confidential, financial or personal. The most sensitive of them are kept on the profile by name.
WHAT IT IS WIRED TO
The MCP servers behind it.
Outgoing edges to MCP servers and tools are listed on the same profile, beside the reach they imply. A tool you have never heard of is a finding, not a footnote.
The attestation field on an exposure profile reads unmeasured. Nothing on this path has been attested yet, and an evidence-bearing field defaults to the unavailable state rather than to a confident one.
[ 05 ]BEFORE YOU DEPLOY IT+
Simulate the rollout,
then let it refuse.
The simulator takes the principals a tool would be given, walks the blast radius of each, and scores what comes back. One of its four verdicts is a refusal to answer, and that is the one that matters most.
FOUR VERDICTS, ONE OF THEM A REFUSAL
insufficient_data
No connected identities, or nothing reachable to evaluate. The score is zero and the only recommendation is to connect a source first. Over an empty graph this is the verdict you get; safe_to_deploy is never returned from zero observed reach.
safe_to_deploy
Least-privilege policies satisfied across what was actually observed. It is a statement about the graph in front of it, not a guarantee about your estate.
conditional_approval
Deployable behind read-only origin constraints, with audit logging configured on retrieval interactions.
blocked_high_risk
Confidential reach, anonymous public links, or both, inside the simulated blast radius. The recommendations name what to revoke before the rollout, in order.
[ 06 ]WHAT THIS PAGE DOES NOT CLAIM+
The gaps, before
somebody else finds them.
The same accounting the security page applies to the authority model, applied to this module.
- No usage analytics. Prompt volumes, seat counts and per-application adoption for third-party AI tools are not collected, because nothing ingests them. There is no chart here that we cannot source.
- No agentless discovery. Ten connector kinds are registered against one fail-closed provider. Until a real upstream client is written for a kind, it is NOT CONNECTED, and it says so in those words.
- Shadow is the default. An agent running under Zorro is not being blocked unless somebody turned enforcement on for that machine. The prompt rate is why, and it is on the record above.
- The kernel view is Linux only. The eBPF and BPF-LSM sensor sees the subprocess a hook cannot, and it reads every kernel offset from the running kernel rather than falling back to a constant. The macOS collector is written against Apple Endpoint Security but cannot run without an Apple-granted entitlement, so it reports UNAVAILABLE instead of guessing. The Windows collector reads the kernel process session and, today, each record's header. Three-platform coverage does not ship.
- No detection rate. The replay harness exists; no benchmark corpus has been run against it. No prevention or detection percentage is claimed here or anywhere else on this site.
- No effector but revocation. A state change on a tool is a record of a decision. Disabling a third-party application, removing its grant or cutting its token would need a connector that is not written, and the product does not pretend the state change did it.
- Scores order a queue. Risk scores and blast-radius scores are ordinals for triage. They are not probabilities, they are not measured, and nothing is auto-blocked on a threshold.
[ 07 ]THE REST OF THE PLATFORM+
One graph,
read four other ways.
The same entities and the same provenance, answering a different question on each page.
START WITH ONE AGENT
Point it at the
agent you trust least.
Run it in shadow, read what it would have decided, and decide whether the answer is worth enforcing.
Build with Zorro ↗