02 / DATA EXPOSURE
Classify the file
without opening it.
Sensitivity from metadata, exposure from the share graph, and an honest answer about what stops working if you revoke. Zero bytes of content are read, and that is an invariant carried in the code, not a promise made in a datasheet.
THE LOOP THIS PRODUCT HAS TO CLOSE
EDGE + GITHUBingest
The only live inbound paths. Ten connector kinds are registered and fail closed.
BUILTgraph
Entities, edges and provenance, projected from events that really happened.
BUILT, THIN INPUTenrich
Classification and entitlement resolution over whatever the graph actually holds.
BUILTanalyze
Exposure scans, posture rules, path search, rollout simulation.
BUILT, NO BASELINEprofile
Per-entity reach and identity profiles. No behavioural or peer baseline.
ONE ACTIONremediate
Propose, approve, execute, verify. The only effector is revoking a session.
WHAT DOES NOT ARRIVE
The assets this module classifies have to come from somewhere, and today no drive, no content platform and no warehouse is connected. Every one of those connector kinds resolves to the same fail-closed provider, so reconciliation over an unfed graph returns no assets and a scan returns no findings. The engines on this page are real. Their input is the break. The ingestion section, in full ↗
WHAT CANNOT BE ACTED ON
There is no revoke-share effector, and there cannot be one without a connector to act through. Resolving an exposure finding does not remove a link, change a permission or contact anyone: it moves the finding to pending re-observation, where it stays until the offending edge is observed to be gone. The one action the control plane can actually execute is revoking an agent session.
This is the module where the break matters most, so it is stated first and again in full further down. Pointed at a graph with no assets in it, this module returns nothing. That is the correct answer and it is the answer you will get.
[ 01 ]RECONCILE+
One asset list,
rebuilt from the graph.
Reconciliation walks every file, dataset and database entity, classifies each from its metadata and writes the result back. The provider comes from the entity's own attribute, or, failing that, from the provenance of whatever observed it.
WHAT ONE ASSET ROW CARRIES
- Asset id and tenant
- Provider, or the observer
- Kind: file, dataset, database
- Path or URL
- MIME type and size
- Sensitivity label
- The rules that matched
- Bytes read, always zero
- Last accessed, if supplied
- First recorded, from provenance
Bytes read is on every row, and it is always zero. It is stored rather than assumed, so the claim on this page is checkable against the data the product produced.
[ 02 ]CLASSIFICATION+
Metadata only.
Zero bytes read.
Classification matches the name, the path and any label the source already applied. It returns the label, the rules that matched, and the number of bytes it read.
SENSITIVITY LABELS
secretspiifinancialipinternalpublic
Every label arrives with its reasons, written as rules you can read back: rule:extension_secret for a key or certificate extension, rule:filename_pattern for a payroll or board-deck name, rule:saas_label when the source already said so, and rule:default_internal when nothing matched. A classification you cannot argue with is a classification you cannot check.
This is pattern matching over names, paths and labels. It will not find a credential inside a file called notes.txt, and it does not claim to. The honest failure mode of metadata-only classification is the false negative, and we would rather carry that one than open your files to avoid it.
[ 03 ]EXPOSURE+
Four ways a file ends up
somewhere nobody chose.
Scanning evaluates each share against the asset behind it. A finding carries the asset, the share, the severity, who it is exposed to, an evidence digest and the time it was detected, and it opens rather than resolving itself.
PUBLIC_LINKAnyone with the link.
A share marked public, or scoped to anonymous. Critical when the asset is labelled secrets, personal or financial; high otherwise. The finding names the share, the URL and what it is exposed to, in that order.
EXTERNAL_DOMAINOutside the boundary.
A share scoped external, or a recipient on a domain that is not yours. Critical when the asset is labelled secrets. The recipient is named, because a finding you cannot act on is a notification.
OVERSHAREDCompany-wide.
A share scoped to the whole company, or named for everyone or all employees. Everybody is not an audience. It is the absence of one, and it is the quietest way a sensitive file becomes reachable by an agent.
STALE_SHARENinety days untouched.
A share that outlived its reason. The threshold is ninety days, fixed in code, and it needs a last-accessed time the source must supply. Without one, no stale finding is raised rather than a guess being made.
Severity escalates on the label, not on a model's opinion: a public link to a secret is critical, a public link to an internal document is high. The rule is small enough to read and fixed enough to argue with.
[ 04 ]WHAT BREAKS+
Before anyone touches it,
say what stops working.
Zorro cannot revoke the share. What it can do is tell the person who will exactly what they are about to break, before they do it rather than after the complaint. It inspects the share's neighbours in the graph and answers in counted terms.
WHAT IT COUNTS
People, services, agents.
Active users on the share, machine identities that depend on it, and AI agents attached to it, each counted from real edges rather than estimated from a name.
WHAT IT CHECKS
Whether there is another way in.
For each affected person it traces an alternative access path. A service account with no second route is the reason a disruption risk reads high, and the reason the first recommendation is to migrate it rather than to revoke.
WHAT IT RETURNS
A risk level and an order.
None, low, medium or high, with an impact summary and the step that has to come first: migrate the service account, add the group, notify the collaborator, or revoke immediately because nothing depends on it.
Be clear about what happens next, because this is where a data-security product usually overstates itself. Resolving an exposure finding here does not revoke the share. There is no revoke-share effector, because there is no connector to act through. The finding moves to pending re-observation and waits for the edge to be observed gone, which is an honest bookkeeping of work somebody still has to do by hand.
[ 05 ]INGESTION, PLAINLY+
What is connected,
and what is not.
This is the fact we would be attacked on, so we would rather publish it ourselves. Agentless ingestion breadth is real, it is valuable, and we do not have it yet. What we do have is the plane underneath it: the decision at the moment an agent acts.
INGESTING TODAY
The edge
IMPLEMENTED
The workstation daemon relays the hook’s log as redacted metadata: decisions, context ingress, lifecycle, guidance deliveries. The projector turns those events into agent, session, activity, secret and network-destination entities, carrying the real event ids as evidence and reporting a cut window rather than presenting a partial answer as complete.
GitHub
IMPLEMENTED
A signature-verified App webhook, normalized into events. A member removal also revokes that reader from the repository’s access list, so the scoping the assistant relies on follows the membership change.
REGISTERED, NOT CONNECTED
Identity providers
Entra ID and Okta are registered connector kinds behind a fail-closed provider. With no credential it answers NOT CONNECTED. With a credential it answers UNAVAILABLE, client not implemented, no upstream call was made. A sync returns not-configured.
Drives and content
Google Workspace, Google Drive, Microsoft 365 and Box: the same provider and the same two answers. Nothing on this site should be read as saying we index your drive today.
Warehouses and workplaces
Snowflake, Salesforce and Slack: the same provider and the same two answers.
What is real here
The framework. The registry, the connector SDK, credential sealing, sync-run recording and tenant isolation are written and tested. What is missing is the upstream client for each kind, and the product refuses to simulate one.
[ 06 ]WHAT THIS PAGE DOES NOT CLAIM+
No number here
that we did not measure.
There are no measured results on this page, and so there are no figures on it.
- No file counts, no reduction figure. No percentage of exposure removed, no count of files discovered, no time saved. The numbers Zorro has measured are all on the security page, and none of them belongs to this module.
- Classification is rules, not a model. No language model, no sampling, no content. The rule list is finite and readable, and it will miss whatever its patterns do not name.
- Stale is a fixed ninety days. Not a learned baseline, not tunable today, and silent when the source supplies no last-accessed time.
- What-breaks is graph truth. It is as good as the edges the graph holds and no better. With no share edges it reports no disruption, which is a statement about the graph rather than about your environment.
- Nothing here revokes a share. The module finds and explains exposures; it cannot remove one. A resolved finding reads pending re-observation until the edge is verified gone, which means the dashboard looks worse than a product that flips a status and calls it done, and means what it says.
[ 07 ]THE REST OF THE PLATFORM+
The same graph,
a different question.
Data exposure is one read of the workforce graph. Here are the others.
BRING THE HARD CASE
Ask what an agent
can already reach.
Bring a repository and an agent, not a drive estate. That is the scope where this answers today, and we will tell you when it is not.
Define your first use case ↗