Introducing Zorro Security / A new perspective on AI security ↗
zorroSECURITYSign in Build with us ↗

02 / DATA EXPOSURE

Classify the file
without opening it.

Sensitivity from metadata, exposure from the share graph, and an honest answer about what stops working if you revoke. Zero bytes of content are read, and that is an invariant carried in the code, not a promise made in a datasheet.

THE LOOP THIS PRODUCT HAS TO CLOSE

EDGE + GITHUB

ingest

The only live inbound paths. Ten connector kinds are registered and fail closed.

BUILT

graph

Entities, edges and provenance, projected from events that really happened.

BUILT, THIN INPUT

enrich

Classification and entitlement resolution over whatever the graph actually holds.

BUILT

analyze

Exposure scans, posture rules, path search, rollout simulation.

BUILT, NO BASELINE

profile

Per-entity reach and identity profiles. No behavioural or peer baseline.

ONE ACTION

remediate

Propose, approve, execute, verify. The only effector is revoking a session.

WHAT DOES NOT ARRIVE

The assets this module classifies have to come from somewhere, and today no drive, no content platform and no warehouse is connected. Every one of those connector kinds resolves to the same fail-closed provider, so reconciliation over an unfed graph returns no assets and a scan returns no findings. The engines on this page are real. Their input is the break. The ingestion section, in full ↗

WHAT CANNOT BE ACTED ON

There is no revoke-share effector, and there cannot be one without a connector to act through. Resolving an exposure finding does not remove a link, change a permission or contact anyone: it moves the finding to pending re-observation, where it stays until the offending edge is observed to be gone. The one action the control plane can actually execute is revoking an agent session.

This is the module where the break matters most, so it is stated first and again in full further down. Pointed at a graph with no assets in it, this module returns nothing. That is the correct answer and it is the answer you will get.

One asset list,
rebuilt from the graph.

Reconciliation walks every file, dataset and database entity, classifies each from its metadata and writes the result back. The provider comes from the entity's own attribute, or, failing that, from the provenance of whatever observed it.

WHAT ONE ASSET ROW CARRIES

  • Asset id and tenant
  • Provider, or the observer
  • Kind: file, dataset, database
  • Path or URL
  • MIME type and size
  • Sensitivity label
  • The rules that matched
  • Bytes read, always zero
  • Last accessed, if supplied
  • First recorded, from provenance

Bytes read is on every row, and it is always zero. It is stored rather than assumed, so the claim on this page is checkable against the data the product produced.

Metadata only.
Zero bytes read.

Classification matches the name, the path and any label the source already applied. It returns the label, the rules that matched, and the number of bytes it read.

SENSITIVITY LABELS

secretspiifinancialipinternalpublic

Every label arrives with its reasons, written as rules you can read back: rule:extension_secret for a key or certificate extension, rule:filename_pattern for a payroll or board-deck name, rule:saas_label when the source already said so, and rule:default_internal when nothing matched. A classification you cannot argue with is a classification you cannot check.

This is pattern matching over names, paths and labels. It will not find a credential inside a file called notes.txt, and it does not claim to. The honest failure mode of metadata-only classification is the false negative, and we would rather carry that one than open your files to avoid it.

Four ways a file ends up
somewhere nobody chose.

Scanning evaluates each share against the asset behind it. A finding carries the asset, the share, the severity, who it is exposed to, an evidence digest and the time it was detected, and it opens rather than resolving itself.

PUBLIC_LINK

Anyone with the link.

A share marked public, or scoped to anonymous. Critical when the asset is labelled secrets, personal or financial; high otherwise. The finding names the share, the URL and what it is exposed to, in that order.

EXTERNAL_DOMAIN

Outside the boundary.

A share scoped external, or a recipient on a domain that is not yours. Critical when the asset is labelled secrets. The recipient is named, because a finding you cannot act on is a notification.

OVERSHARED

Company-wide.

A share scoped to the whole company, or named for everyone or all employees. Everybody is not an audience. It is the absence of one, and it is the quietest way a sensitive file becomes reachable by an agent.

STALE_SHARE

Ninety days untouched.

A share that outlived its reason. The threshold is ninety days, fixed in code, and it needs a last-accessed time the source must supply. Without one, no stale finding is raised rather than a guess being made.

Severity escalates on the label, not on a model's opinion: a public link to a secret is critical, a public link to an internal document is high. The rule is small enough to read and fixed enough to argue with.

Before anyone touches it,
say what stops working.

Zorro cannot revoke the share. What it can do is tell the person who will exactly what they are about to break, before they do it rather than after the complaint. It inspects the share's neighbours in the graph and answers in counted terms.

WHAT IT COUNTS

People, services, agents.

Active users on the share, machine identities that depend on it, and AI agents attached to it, each counted from real edges rather than estimated from a name.

WHAT IT CHECKS

Whether there is another way in.

For each affected person it traces an alternative access path. A service account with no second route is the reason a disruption risk reads high, and the reason the first recommendation is to migrate it rather than to revoke.

WHAT IT RETURNS

A risk level and an order.

None, low, medium or high, with an impact summary and the step that has to come first: migrate the service account, add the group, notify the collaborator, or revoke immediately because nothing depends on it.

Be clear about what happens next, because this is where a data-security product usually overstates itself. Resolving an exposure finding here does not revoke the share. There is no revoke-share effector, because there is no connector to act through. The finding moves to pending re-observation and waits for the edge to be observed gone, which is an honest bookkeeping of work somebody still has to do by hand.

What is connected,
and what is not.

This is the fact we would be attacked on, so we would rather publish it ourselves. Agentless ingestion breadth is real, it is valuable, and we do not have it yet. What we do have is the plane underneath it: the decision at the moment an agent acts.

INGESTING TODAY

The edge
IMPLEMENTED
The workstation daemon relays the hook’s log as redacted metadata: decisions, context ingress, lifecycle, guidance deliveries. The projector turns those events into agent, session, activity, secret and network-destination entities, carrying the real event ids as evidence and reporting a cut window rather than presenting a partial answer as complete.
GitHub
IMPLEMENTED
A signature-verified App webhook, normalized into events. A member removal also revokes that reader from the repository’s access list, so the scoping the assistant relies on follows the membership change.

REGISTERED, NOT CONNECTED

Identity providers
Entra ID and Okta are registered connector kinds behind a fail-closed provider. With no credential it answers NOT CONNECTED. With a credential it answers UNAVAILABLE, client not implemented, no upstream call was made. A sync returns not-configured.
Drives and content
Google Workspace, Google Drive, Microsoft 365 and Box: the same provider and the same two answers. Nothing on this site should be read as saying we index your drive today.
Warehouses and workplaces
Snowflake, Salesforce and Slack: the same provider and the same two answers.
What is real here
The framework. The registry, the connector SDK, credential sealing, sync-run recording and tenant isolation are written and tested. What is missing is the upstream client for each kind, and the product refuses to simulate one.

No number here
that we did not measure.

There are no measured results on this page, and so there are no figures on it.

The same graph,
a different question.

Data exposure is one read of the workforce graph. Here are the others.

BRING THE HARD CASE

Ask what an agent
can already reach.

Bring a repository and an agent, not a drive estate. That is the scope where this answers today, and we will tell you when it is not.

Define your first use case ↗