Introducing Zorro Security / A new perspective on AI security ↗
zorroSECURITYSign in Build with us ↗

01 / AI & AGENT SECURITY

Know every agent.
Decide before it acts.

An inventory tells you which agents exist. It does not tell you what they reach through permissions they borrowed, and it cannot stop the next tool call. This module does both, and reports the empty case as empty.

THE LOOP THIS PRODUCT HAS TO CLOSE

EDGE + GITHUB

ingest

The only live inbound paths. Ten connector kinds are registered and fail closed.

BUILT

graph

Entities, edges and provenance, projected from events that really happened.

BUILT, THIN INPUT

enrich

Classification and entitlement resolution over whatever the graph actually holds.

BUILT

analyze

Exposure scans, posture rules, path search, rollout simulation.

BUILT, NO BASELINE

profile

Per-entity reach and identity profiles. No behavioural or peer baseline.

ONE ACTION

remediate

Propose, approve, execute, verify. The only effector is revoking a session.

WHAT DOES NOT ARRIVE

Nothing discovers third-party AI applications on its own. The inventory is whatever the edge and GitHub put into the graph: agents, sessions, the authority decisions they produced, and the secrets and destinations a finding named. Identity, drive and SaaS connectors are registered against a fail-closed provider that answers NOT CONNECTED without a credential and UNAVAILABLE with one, so there is no usage analytics and no vendor-side discovery behind any of this. What is connected today ↗

WHAT CANNOT BE ACTED ON

The control plane has exactly one effector: revoking an agent session, which takes authority away and can never grant it. Marking a tool revoked changes our record of it. It does not reach a vendor, disable an account or remove anybody’s access, because there is no connector to act through. The enforcement that does exist is the hook, on the machine, on the next tool call.

A module is a drawing; a loop has to close. Zorro owns the middle of this one and is missing both ends, so read every capability below against the two cells above it. Discovery is exactly as populated as the graph is, and an unpopulated graph returns an empty inventory rather than a reassuring one.

Four kinds of thing
that act on your behalf.

Discovery walks the graph for AI tools and agents, keeps whatever a person has already reviewed, and classifies only what is new. Nothing is overwritten by a rediscovery.

ENTITY KINDS THIS MODULE READS

ai_agentai_toolmcp_servermcp_toolaccountsessionsecretnetwork_destination

Every discovered entity keeps the provenance it arrived with: the source that observed it, the evidence_digest behind that observation, and the time it was recorded. An entry with no provenance is not an entry we are willing to show you.

The inventory is the map.
The hook is the gate.

One binary reads a hook payload on standard input, answers on standard output and exits. No network, no daemon to reach, no lock to take, once per tool call, on the agent's critical path.

FIVE HARNESSES

One answer format each.

Claude Code, Codex, Cursor, GitHub Copilot and Gemini CLI each send their own payload and expect their own reply. Which one is calling comes from the flag, then the environment, then the shape of the payload itself.

SHADOW FIRST

The default blocks nothing.

The first run on any machine records what it would have decided. That is how the origin sets get derived from what a repository actually does, and how the prompt rate is measured before anyone is interrupted. Enforcement is opt-in.

FAILURE ASKS

Never a silent allow.

A payload that will not parse, a ledger that will not open, a tool or a provider that is not recognised: every one of them becomes a question for a person. The absence of an answer is not an answer.

ONE REFUSAL

A revoked session.

A person, or the control plane through the workstation daemon, stopped it. Nothing that session proposes proceeds, and the refusal lands on the next tool call rather than in a report.

THE MCP PROXY

caveat-mcp

on the stdio relay, before the harness sees the answer

Every tool listing is reduced to a pin per tool, a hash over the canonical JSON of the whole tool object plus the server's identity, recorded before the answer is passed on. The first observation is the baseline, so the hook can never see a definition the store has not. Tool descriptions, initialize instructions, results and sampling requests are all recorded as untrusted content, which means they can only narrow a live grant. The server also runs on an allowlisted environment rather than the harness's own: tokens in the parent are not the server's to read.

WHAT THE DECISION COSTS

caveat-bench, on recorded sessions

measured, with the caveat attached

111 to 640 µs p50, with p99 in milliseconds and machine-load dependent. The number we are not happy with is the other one: 36.66 mean and 25 p50 prompts per session on a frozen transcript set, against a target of two or fewer. That is the active engineering priority, and it is the honest reason shadow is the default rather than a setting we hope you change.

Sanctioned, shadow,
or honestly under review.

The posture the entity carries decides the state. Nothing infers approval from popularity, from usage volume or from a vendor's reputation.

FOUR STATES, RECORDED ON THE TOOL

under_review

The default for anything newly discovered. It is not a verdict, it is the absence of one, and it is recorded as such rather than defaulting to approved.

sanctioned

Approved by a named person. The review writes the reviewer, the time and the note onto the tool row, so the approval has an author.

shadow

Present and not approved. The exposure profile raises it as a finding: an unsanctioned tool connected without administrative review.

revoked

Approval withdrawn in the record. The row stays, because deleting it would delete the history of having approved it. Withdrawing approval is a decision, not an enforcement: nothing here reaches a vendor and turns the tool off.

The risk score attached to each state is an ordinal label for sorting a queue. It is set by the state, not derived from observed behaviour, and it is not a probability. Read the state; the number is only there to order the list.

The permission it borrowed
is the one that matters.

An agent rarely holds access of its own. It runs as somebody, and it reaches whatever that somebody reaches, which is why an agent inventory that stops at the agent is not an answer.

WHO IT RUNS AS

Operating principals, read from edges.

Incoming uses and authenticates_as edges name the identities an agent acts through. It is a graph read, not an assumption about what a token probably carries.

WHAT IT REACHES

Traced, then counted by sensitivity.

The access path is walked to the files, data sources and databases at the end of it, and what is found is counted as confidential, financial or personal. The most sensitive of them are kept on the profile by name.

WHAT IT IS WIRED TO

The MCP servers behind it.

Outgoing edges to MCP servers and tools are listed on the same profile, beside the reach they imply. A tool you have never heard of is a finding, not a footnote.

The attestation field on an exposure profile reads unmeasured. Nothing on this path has been attested yet, and an evidence-bearing field defaults to the unavailable state rather than to a confident one.

Simulate the rollout,
then let it refuse.

The simulator takes the principals a tool would be given, walks the blast radius of each, and scores what comes back. One of its four verdicts is a refusal to answer, and that is the one that matters most.

FOUR VERDICTS, ONE OF THEM A REFUSAL

insufficient_data

No connected identities, or nothing reachable to evaluate. The score is zero and the only recommendation is to connect a source first. Over an empty graph this is the verdict you get; safe_to_deploy is never returned from zero observed reach.

safe_to_deploy

Least-privilege policies satisfied across what was actually observed. It is a statement about the graph in front of it, not a guarantee about your estate.

conditional_approval

Deployable behind read-only origin constraints, with audit logging configured on retrieval interactions.

blocked_high_risk

Confidential reach, anonymous public links, or both, inside the simulated blast radius. The recommendations name what to revoke before the rollout, in order.

The gaps, before
somebody else finds them.

The same accounting the security page applies to the authority model, applied to this module.

One graph,
read four other ways.

The same entities and the same provenance, answering a different question on each page.

START WITH ONE AGENT

Point it at the
agent you trust least.

Run it in shadow, read what it would have decided, and decide whether the answer is worth enforcing.

Build with Zorro ↗