Classification, DLP and insider risk

Know where your sensitive data is, where it moves, and who is about to take it

Guardian Sentinel classifies data at rest and in motion, maps every flow between the systems that hold it, and enforces one policy across endpoint, browser, email, SaaS and cloud storage. When an identity starts behaving like an exit interview, the platform sees it before the upload finishes.

240+ Built-in classifiers covering 61 regulatory regimes out of the box
18ms Median inline inspection latency added to a file write or upload
99.2% Classification precision on structured identifiers after tuning
-74% Fewer DLP false positives after 30 days of context-aware scoring
7yr Maximum retention for data-movement evidence on the Enterprise plan

The approach

Legacy DLP failed because it only ever saw the last three feet

Pattern-matching at the egress point produces two outcomes: it blocks the invoice your finance team needed to send, and it misses the source code pasted into a chat window. Sentinel decides using content, lineage, identity risk and destination reputation together — four inputs instead of one regular expression.

Classify once, at the source

Content is fingerprinted where it is created and the label travels with it — through copies, renames, archives, screenshots and format conversions. A file does not become unclassified because someone zipped it.

Judge with lineage, not just content

A spreadsheet of 400 numbers is ambiguous. A spreadsheet of 400 numbers exported from the production customer database nine minutes ago by an account under notice is not.

Enforce with proportion

Five graded responses — allow, log, warn and justify, encrypt in place, block — chosen by the same policy engine, so the business keeps moving and the exception is the exception.

Classification

Four detection methods, one label

Regular expressions alone are not a classification strategy. Sentinel combines pattern matching, exact data matching against your own records, document fingerprinting and a trained content model, then reconciles them into a single confidence-weighted label.

INPUT DETECTORS RECONCILIATION POLICY OUTCOME Content event file write · upload message send · paste database export · print Pattern + checksum validator Exact data match index Document fingerprint Trained content model + OCR on device · 4 ms · 240 classifiers salted hash · 41M records indexed rolling shingles · survives edits local inference · 12 ms · no upload Label reconciliation highest-confidence wins, agreement raises confidence restricted:pci 0.92 + lineage · + identity risk · + destination trust Allow and log Warn and justify Encrypt in place Redact and release Block and capture evidence Nothing in this pipeline requires content to leave the device or tenant — only the verdict and its metadata are published.

Scroll horizontally to see the full diagram on a small screen.

How each method behaves, and where Sentinel applies it by default.
Method Best for Precision Where it runs
Pattern and validator Card numbers, national IDs, IBANs, access keys — anything with a checksum 99.2% Endpoint, gateway, cloud
Exact data match Your actual customer, patient or employee records, hashed and indexed 99.9% Gateway, cloud, endpoint (partial index)
Document fingerprint Contracts, board packs, designs, source files and their derivatives 98.6% Endpoint, cloud, email
Trained content model Unstructured categories — legal advice, M&A material, clinical notes 94.1% Cloud, email, browser
Optical character recognition Screenshots, scanned documents and images pasted into chat 96.3% Endpoint, browser, email
Inherited label Microsoft Purview, Titus and Boldon James labels already applied 100% Everywhere the label is readable

Coverage that maps to your obligations

Classifier packs ship pre-mapped to PCI DSS 4.0, HIPAA, GDPR, CCPA, DORA, NIS2, GLBA, ITAR, and 53 further regimes. Each pack declares exactly which data elements it detects and which control it satisfies, so audit evidence writes itself.

Your own classifiers in minutes

Build custom classifiers from a sample set, a regular expression with a validation function, or a structured extract from a database. New classifiers are backtested against 30 days of historical content before you enable them, so you see the false-positive rate before your users do.

Exact data matching keeps your records private Source records are salted and hashed on your infrastructure. Only the index leaves; the underlying values never do.

Data lineage

Every copy of your crown jewels, and every route out

Sentinel builds a live flow map from actual observed movement — not from a data-flow diagram someone drew in 2021. Volumes are measured, destinations are reputation-scored, and any edge can be turned into a policy in one click.

SOURCES OF RECORD PROCESSING AND HANDLING DESTINATIONS Customer database PostgreSQL · 41M records Restricted · PCI + GDPR Salesforce tenant CRM · 2.6M opportunity records Confidential Document repository SharePoint · 8.4TB Mixed · 12% restricted Analytics warehouse tokenised at load · 4.1 GB/day Policy: allow, log Managed endpoints 11,900 devices · 780 MB/day Policy: warn and justify Email gateway 2.4M messages/day Policy: encrypt if restricted Partner SFTP · sanctioned contractual flow · 3.9 GB/day · encrypted Managed cloud backup customer-managed keys · 1.2 GB/day Personal cloud storage BLOCKED · 214 attempts · 180 MB Unmanaged USB device BLOCKED · 37 attempts · 6.2 GB Sanctioned flow Requires justification Blocked exfiltration attempt Band thickness is proportional to 24-hour volume.

Scroll horizontally to see the full map on a small screen.

Anonymised flow map from a financial services tenant, 72 hours after deployment. Two of the four destinations were unknown to the security team.

Enforcement

One policy, seven control points, no rewriting

Write the rule once in plain policy language. Sentinel compiles it to every control point that can enforce it and tells you honestly where enforcement is degraded — because a silent gap is worse than a documented one.

Enforcement capability by control point, release 2026.7.
Control point Inspect Warn user Block Encrypt or redact Typical added latency
Endpoint agent (Windows, macOS, Linux) Yes Yes Yes Yes 18 ms
Browser extension (Chrome, Edge, Firefox) Yes Yes Yes Redact only 30 ms
Email — Microsoft 365 and Google Workspace Yes Yes Yes Yes 1.2 s
SaaS API connectors (out of band) Yes Yes Quarantine Yes 40 s
Cloud object storage (S3, Blob, GCS) Yes n/a Yes Yes Near real time
Network egress via Guardian or partner proxy Yes Yes Yes No 9 ms
Removable media and printing Yes Yes Yes Yes 22 ms

Scroll horizontally to see all columns.

Policy that reads like an instruction

Policies are authored in a declarative language, stored in Git, reviewed as code and promoted through environments. Nothing is clicked into a console and forgotten.

policy "restricted-customer-data-egress" {
  match {
    classification = ["restricted:pci", "restricted:pii"]
    min_confidence = 0.85
  }
  when destination.trust != "sanctioned" {
    action  = "block"
    notify  = ["user", "manager", "soc"]
    capture = "evidence_bundle"
  }
  when identity.risk >= 60 {
    action = "block"
  }
  otherwise {
    action = "warn_and_justify"
    ttl    = "8h"
  }
}

Five graded outcomes, not two

  • Allow and log — the flow is sanctioned; a record is kept for lineage, audit and retrospective hunting.
  • Warn and justify — the user sees a contextual prompt naming the specific data, and their business reason is recorded against the event.
  • Encrypt in place — the file or message is protected with a policy key so it remains readable only to authorised recipients wherever it travels.
  • Redact and release — matched elements are masked and the remainder is delivered, which keeps most legitimate workflows unblocked.
  • Block and capture — the transfer is stopped and a forensic evidence bundle is preserved, hashed and retained for legal hold.
Honest coverage reporting If a device is offline, an extension is missing or a SaaS tenant grants read-only API scope, Sentinel marks that path as degraded on the coverage dashboard instead of implying protection it cannot deliver.

Exfiltration prevention

The staging is the signal. The upload is already too late.

Bulk theft has a shape: collect, compress, encrypt, split, then move. Sentinel scores that sequence as it unfolds, which is why most interventions happen before a single byte reaches the internet.

Signals that precede the transfer

  • Recursive read of a file share at a rate no human interaction pattern produces
  • Archive creation with an unusually high compression ratio and password protection
  • Volume-shadow copy access or database export outside a scheduled job window
  • Splitting into uniform chunks sized to evade an attachment or upload threshold
  • DNS or ICMP tunnelling, and encrypted channels to newly registered domains
  • Cloud storage sync client installed within the last hour on a managed device
  • Print-to-PDF of a hundred records followed by a personal webmail session

A real interception, in sequence

  1. 23:41:06 · T+0

    Bulk read begins

    A contractor account starts a recursive read across a legal file share, touching 6,400 documents in 90 seconds. Sentinel raises a data-staging signal at confidence 0.62.

  2. 23:43:20 · T+2m 14s

    Encrypted archive created

    A 2.1 GB password-protected archive is written to a temporary directory. The lineage engine links it to 4,180 documents classified restricted. Confidence rises to 0.88.

  3. 23:44:02 · T+2m 56s

    Sync client installed

    A consumer cloud storage client is installed and authenticated with a personal account. Destination trust is unsanctioned. Policy threshold crossed.

  4. 23:44:03 · T+2m 57s

    Blocked and contained

    Upload blocked at the endpoint, the archive quarantined and hashed, the account's sessions revoked, and the device isolated. Zero bytes of restricted content left the estate.

  5. 23:47:15 · T+6m 09s

    Case delivered with evidence

    The SOC received a single case containing the document inventory, the archive hash, the destination account identifier and a chain-of-custody record admissible for HR and legal proceedings.

Insider risk

Most insider incidents are careless. A few are not. Tell them apart.

Sentinel scores insider risk from behaviour, sequence and context — never from content of personal communications. Scores decay, they are explainable line by line, and the whole model runs under access controls your works council can review.

31% Of data-loss cases involve a departing employee within 30 days of notice Guardian Labs incident data, 2025
3.6x Higher exfiltration attempt rate among unmanaged contractor devices Measured across 900 tenants
68% Of policy violations resolved by the warn-and-justify prompt alone No analyst involvement required
11d Median lead time between first risk signal and an attempted bulk transfer Time you get back

Behavioural signals

Weighted 45

Deviation from an identity's own 30-day baseline: access breadth, working hours, repository cloning, mass downloads, unusual application launches and first-time access to systems outside their role.

Sequence signals

Weighted 35

Ordered patterns that only matter in combination — collect, then compress, then install a sync client. A single step is noise; the sequence is the detection.

Context signals

Weighted 20

HR-provided lifecycle events such as notice period, role change or contract end, consumed as a boolean flag through your HRIS connector. No personal detail is ingested.

coach at 60 contain at 85 0 25 60 85 100 day 14 · repository cloning day 19 · out-of-hours bulk download day 23 · encrypted archive created day 27 · upload blocked, session revoked day 1 day 8 day 16 day 24 day 30

Scroll horizontally to see the full chart on a small screen.

Identities remain pseudonymised until a dual-approval unmasking is granted. The eleven-day lead time between the first signal and the transfer attempt is typical.
Privacy controls that survive a works council review Pseudonymisation by default, so analysts see USR-4A19 until a dual-approval unmasking is granted and logged. Regional data residency, configurable exclusion of protected categories, and a complete audit of every unmasking request.
Scores decay and explain themselves Every score shows its contributing events with weights and timestamps. Signals age out on a configurable half-life, so a single bad week does not follow someone for a year.

Encryption and key management

Protection that stays attached to the file

Sentinel applies policy-based encryption at the moment of classification so a restricted document remains protected after it leaves your estate — on a partner's laptop, in a personal inbox, or on a USB stick found in a car park.

Standards, not proprietary wrappers

AES-256-GCM for content, RSA-4096 or ECDH P-384 for key exchange, and FIPS 140-3 validated modules on every platform. Protected files open in native applications through the Guardian handler or the Microsoft Purview client.

Your keys, your custody

Bring your own key or hold your own key through AWS KMS, Azure Key Vault Managed HSM, Google Cloud KMS, HashiCorp Vault, Thales and Entrust HSMs. Revoke a key and every copy of the data becomes unreadable, everywhere, immediately.

Access that can be withdrawn

Grant time-boxed access to external recipients, watch who opened what and from where, and revoke it after the deal closes. Every open, print and forward attempt is logged against the document, not the mailbox.

Which data protection capability satisfies which control. Evidence exports are generated continuously and are accepted as-is by most assessors.
Regime Control Guardian capability Evidence produced
PCI DSS 4.0 3.3, 3.5, 12.10 PAN discovery, truncation enforcement, egress blocking Cardholder data inventory with location and owner
GDPR Art. 30, 32, 33 Processing inventory, encryption, breach timeline reconstruction Records of processing plus 72-hour notification pack
HIPAA 164.312(a), (b), (e) PHI classification, access logging, transmission security ePHI access report by system and workforce member
DORA Art. 9, 17, 19 Data resilience controls, incident classification, reporting Major incident report with impact and root cause
ISO 27001:2022 A.5.12, A.8.10, A.8.12 Classification scheme, deletion evidence, leakage prevention Statement of applicability annex with live control status
NIS2 Art. 21(2)(a), (d) Risk analysis, supply-chain data flow visibility Third-party data flow register with volumes

Scroll horizontally to see all columns.

Working with an assessor now? Guardian compliance advisory maps your existing controls before deployment.

Deployment

From observe to enforce without a business revolt

Every failed DLP programme enforced on day one. Guardian deployments run in monitor mode until the data says the policy is right.

  1. Discover

    Connect read-only to storage, SaaS and email. Within 48 hours you have a classified inventory, a flow map and a list of destinations nobody approved. No enforcement, no user impact.

  2. Model

    Author policies against real observed traffic and replay them across 30 days of history. The platform reports precisely how many transfers each rule would have blocked, and which teams they belonged to.

  3. Coach

    Enable warn-and-justify first. Roughly two thirds of violations stop here, users learn the rule in context, and the security team gets a clean signal of what is genuinely business-critical.

  4. Enforce

    Turn on blocking for the highest-confidence classifications and unsanctioned destinations, one business unit at a time, with an exception workflow that resolves in minutes rather than weeks.

Our previous DLP generated 900 events a day and we actioned maybe four of them. Guardian gave us 30 a day that all meant something, and it caught a departing engineer staging our design files three weeks before his last day.

Daniel Okafor
Head of Information Security, medical device manufacturer

Questions

What data teams ask us first

Does inline inspection slow down our users?

Median added latency is 18 milliseconds on endpoint file operations and 30 milliseconds in the browser extension. Inspection runs locally against a compiled classifier set, so no content leaves the device to reach a decision. Files above a configurable size are inspected asynchronously with the transfer held, which keeps large media workflows usable.

Can Guardian read our sensitive content?

No. Classification happens on your infrastructure or in your tenant region. What reaches the Guardian control plane is the verdict, the classifier that matched, a confidence value and metadata. Optional evidence capture stores a snippet only where policy requires it, encrypted with your key, with a configurable retention window and an access log.

How does this coexist with Microsoft Purview?

Sentinel reads and writes Purview sensitivity labels, so existing labels are honoured and new classifications are published back. Most customers keep Purview for labelling and document protection inside Microsoft 365, and use Guardian for cross-platform enforcement, lineage and insider risk — the parts Purview does not reach.

What happens when a device is offline?

The endpoint agent caches a compiled policy and the classifier set, so inspection and enforcement continue with no connectivity. Events queue locally in an encrypted store and synchronise on reconnection, preserving original timestamps. Offline policy has a configurable maximum age after which the agent fails closed for restricted classifications.

Can we scope insider risk to specific populations?

Yes, and most customers must. Scope by legal entity, country, department or employment type; exclude protected categories entirely; and require dual approval before any identity is unmasked. All configuration is exportable as a document you can hand to a works council, regulator or data protection officer.

How long does a first deployment take?

Discovery is live within a day for SaaS and cloud storage, and within a week for endpoints at typical enterprise scale. Most customers reach enforced blocking on their top three data classes within six weeks. Guardian consulting runs the policy design workshops if you would rather not build the first ruleset alone.

See your real data flow map in 48 hours

Connect one SaaS tenant and one storage account, read-only. Guardian returns a classified inventory, a live flow map and the list of destinations your data is already reaching. No agents, no enforcement, no obligation.