Security orchestration, automation and response

Containment at machine speed, with a human in exactly the right place

Breakout time is now 41 minutes. No analyst rota wins that race by reading alerts. Sentinel investigates, decides and contains autonomously within a boundary you define — and escalates to a person only where judgement genuinely changes the outcome.

-93% Median reduction in mean time to respond after 90 days
4.1s Median time from detection to enforced containment action
96% Of alerts triaged to a verdict without an analyst touching them
0.4% Autonomous action reversal rate across the customer base
380+ Maintained playbooks and 1,200 typed actions in the response library

The problem with first-generation SOAR

You bought orchestration and inherited a software project

Classic SOAR gave you a canvas and a thousand connectors, then asked your three most senior analysts to become integration engineers. Two years later the playbooks are brittle, the API tokens have rotated, half the branches were never tested, and the automation everyone was promised runs on four use cases.

Sentinel inverts it. Detection, enrichment, decision and action live in one platform with one data model, so a playbook does not spend eleven steps gathering context the platform already holds. What you author is policy — not plumbing.

Context is already local

Asset criticality, identity risk, data classification and intel reputation are resolved in the same query. No 40-second round trip to four external APIs before a decision can be made.

Reasoning, not branching

The Sentinel AI Analyst writes the investigation narrative, states its confidence and cites every artefact it used. Playbooks handle the deterministic parts; judgement is modelled, not hard-coded into 200 if-statements.

Blast radius is bounded

Every action declares its scope, reversibility and maximum affected assets. Guardrails refuse an action that would exceed the limits you set, no matter how confident the detection is.

Everything is reversible

Isolation, disablement, quarantine and rotation all have a tested undo path and a one-click rollback that restores prior state with a full audit record.

Autonomous response tiers

Autonomy is a dial, and you hold it

Most organisations do not want a binary choice between full automation and none. Sentinel expresses autonomy as four tiers that you assign per detection class, per asset group and per time window — so a production database at 03:00 can behave differently from a laptop at noon.

Tier 0

Observe

The playbook runs end to end in shadow mode. Every action it would have taken is recorded with its inputs and confidence, and nothing touches production. This is how every new playbook starts.

Actions taken: none
Tier 1

Assist

Sentinel enriches, correlates and writes the investigation, then presents a recommended action with a single confirm button. Analysts keep authority; the eleven minutes of gathering context disappear.

Read-only automation
Tier 2

Supervised

Reversible containment executes immediately; anything irreversible or wide-scope waits for approval. Isolate the host now, ask before you disable the shared service account.

Act, then notify
Tier 3

Autonomous

Full containment and remediation without human involvement, within declared guardrails. Reserved for high-confidence detections where the cost of waiting exceeds the cost of a rare rollback.

Median 4.1 s to contain
A representative tier assignment. Every row is configurable, and the defaults ship deliberately conservative.
Detection class Confidence floor Managed endpoint Production server Domain controller Executive device
Ransomware encryption behaviour 0.90 Tier 3 Tier 3 Tier 2 Tier 3
Known-malicious binary execution 0.95 Tier 3 Tier 3 Tier 2 Tier 3
Credential dumping (LSASS access) 0.88 Tier 3 Tier 2 Tier 2 Tier 2
Adversary-in-the-middle session theft 0.85 Tier 3 Tier 3 Tier 3 Tier 2
Data staging and bulk exfiltration 0.82 Tier 3 Tier 2 Tier 2 Tier 2
Suspicious cloud role assignment 0.80 n/a Tier 2 Tier 1 n/a
Phishing email with weaponised link 0.75 Tier 3 n/a n/a Tier 3
Anomalous internal reconnaissance 0.70 Tier 2 Tier 1 Tier 1 Tier 1

Scroll horizontally to see all columns.

Tiers can also be scheduled — many customers run Tier 3 outside business hours and Tier 2 during the change freeze window.

Anatomy of a playbook

What actually happens in the four seconds after a detection

This is the shipped ransomware containment playbook, unedited. Timings are the p50 measured across the fleet. Every node is inspectable, versioned in Git, and testable against recorded incidents before it goes anywhere near production.

T+0 ms T+300 ms T+340 ms T+4.1 s T+38 s Trigger mass file entropy change Asset criticality Identity risk score Intel reputation tier 1 finance server 78 / 100 · elevated CORVUS WOLF tooling match Decision confidence 0.97 ≥ 0.90 scope 1 host ≤ limit 25 below floor → approval Contain kill process tree isolate host (agent stays up) revoke identity tokens block C2 at egress suspend shadow copies deletion snapshot volume for forensics Roll back restore 12,401 files from local shadow store Approval gate pages on-call in Slack + PagerDuty 15 min timeout → fail safe Verify re-scan host, confirm no persistence remains Close with evidence narrative, timeline, hashes, actions taken, rollback plan ticket synced to ServiceNow Automated action Decision node Human approval Verification Dashed path is taken only when confidence falls below the configured floor.

Scroll horizontally to see the full flow on a small screen.

Executed 11,904 times across the customer base in the trailing 12 months, with 47 rollbacks — a 0.4% reversal rate.

Playbook library

380 playbooks that Guardian maintains, not you

Every playbook ships versioned, tested against recorded incidents, and updated when adversary tradecraft changes. Fork any of them, or write your own against the same typed action library.

Ransomware containment

Detects encryption behaviour, kills the process tree, isolates the host, blocks the C2 destination and restores affected files from the local shadow store. Median containment 4.1 seconds.

v14.2 · tier 3 · 11,904 runs

Compromised identity response

Revokes tokens, purges Kerberos tickets, disables the account, removes privileged group membership, rotates dependent secrets and reverts unauthorised directory changes.

v9.6 · tier 2 · 6,320 runs

Phishing campaign sweep

Retracts the message from every mailbox that received it, blocks the sender and infrastructure, detonates attachments, and checks whether anyone authenticated to the harvested credential page.

v11.0 · tier 3 · 24,780 runs

Cloud misconfiguration rollback

Reverts a public bucket, an over-permissive security group or a dangerous role trust policy to its last compliant state, and opens a pull request against the offending Terraform module.

v7.3 · tier 2 · 3,915 runs

Data exfiltration interception

Blocks the transfer, quarantines and hashes the staged archive, revokes the session, isolates the device and assembles a chain-of-custody evidence pack for HR and legal.

v6.1 · tier 2 · 1,806 runs

Lateral movement interruption

Applies dynamic micro-segmentation around the affected segment, terminates the offending sessions, and blocks the specific service and port pair rather than quarantining a whole VLAN.

v8.4 · tier 2 · 2,447 runs

Intel-driven retro hunt

When Guardian Labs publishes new indicators, this playbook sweeps retained telemetry, opens cases for historical matches and pre-blocks the infrastructure across every enforcement point.

v5.9 · tier 3 · 8,102 runs

Regulatory notification pack

Assembles the material facts, affected record counts, timeline and control status into a draft filing for GDPR, DORA, NIS2, HIPAA or SEC disclosure within the statutory clock.

v4.2 · tier 1 · 380 runs

Vulnerable asset shielding

On publication of an exploited-in-the-wild CVE, identifies affected assets, applies virtual patching at the sensor, and raises change tickets ranked by real exposure rather than CVSS alone.

v10.1 · tier 2 · 5,633 runs

Approval workflows

Human judgement, delivered where the human already is

An approval request that requires logging into a console at 03:00 is an approval that times out. Sentinel delivers the decision to Slack, Teams, PagerDuty, email or the mobile app with the full context attached and the action pre-authorised.

Approval required INC-88214 · expires in 14:31

Disable service account svc_sqlbackup across 3 domains

Why: credential theft observed on WKS-4471 at 09:16, then this account authenticated to DC01 for the first time in 1,340 days. Confidence 0.91.

Blast radius: 4 dependent services identified — nightly SQL backup, reporting extract, two scheduled tasks. Next scheduled run in 6h 20m.

Reversible: yes, one click, estimated 8 seconds to restore.

Approve Approve & suppress dependents Decline with reason

Illustration of the approval card as it appears in Slack. The controls above are shown for context and are not interactive on this page.

Timeouts fail safe, not fail open

You choose what happens when nobody answers: escalate to the next rota tier, execute the reversible subset, or stand down and keep monitoring. The default is escalate, then execute the reversible subset at 15 minutes.

Dual control for irreversible acts

Actions marked destructive — wiping a device, deleting a cloud resource, forcing a tenant-wide password reset — can require two distinct approvers from separate groups, enforced by the platform.

Where approvals arrive

  • Slack and Microsoft Teams — interactive cards with full context, scoped to the channel that owns the asset.
  • PagerDuty and Opsgenie — approval as a first-class incident action with rota-aware routing and escalation policies.
  • ServiceNow — approval bound to a change record so CAB governance is satisfied without a parallel process.
  • Guardian mobile — biometric-confirmed approvals with the same evidence bundle, for the 3 a.m. case.
  • API and webhook — route approvals into your own workflow engine and post the decision back with a signed callback.
Every approval is evidence Who approved, from which device, on what basis, and what the state was before and after — recorded immutably and exportable for audit, regulator or litigation hold.

Measured outcomes

Where the hours actually go, before and after

MTTR is not one number; it is five stages with very different bottlenecks. Automating triage without automating containment moves the queue, not the risk. Here is the full breakdown from 62 enterprise deployments measured 30 days before and 90 days after.

0 30 60 90 120 median minutes per stage Detection Triage Investigation Containment Recovery & closure 34 min 52 min 96 min 71 min 128 min 24 s 1.2 min 6 min 4.1 s 21 min Before: 381 min MTTR After: 28.7 min · -93% Before Guardian After 90 days of graded autonomy

Scroll horizontally to see the full chart on a small screen.

Medians across 62 enterprise deployments, 2025. Recovery remains the slowest stage because it is the one that still legitimately needs people.
96% Alerts closed to a verdict with no analyst interaction
11h Analyst hours returned per week, per 1,000 endpoints
-68% Reduction in escalations to tier-3 engineering
2.4x More incidents fully closed per analyst per shift

Guardrails

The controls that make autonomy defensible

Autonomous response only survives a board conversation if you can explain precisely what it will never do. These limits are enforced by the platform, not by playbook discipline.

  1. Scope ceilings

    Every action declares a maximum affected asset count and a maximum rate. A playbook that would isolate more than 25 hosts in five minutes halts and requests approval, regardless of confidence. Ceilings are set per playbook, per environment and per time window.

  2. Protected asset lists

    Domain controllers, certificate authorities, safety-critical OT, medical devices and any asset you designate can be marked never-isolate or never-disable. Sentinel will still detect, investigate and alert — it simply will not act unilaterally.

  3. Reversibility classification

    Actions are typed as reversible, reversible-with-effort or destructive. Autonomy tiers map to these types, so no configuration mistake can grant unattended authority over a destructive action.

  4. Change-freeze awareness

    Sentinel reads your change calendar from ServiceNow or Jira. During a declared freeze, playbooks automatically downgrade one tier and route to the on-call approver instead of executing.

  5. Dry run and replay

    Every playbook version can be replayed against 90 days of recorded incidents before promotion. The diff reports exactly which actions would have changed, on which assets, with what outcome.

  6. Immutable audit and one-click rollback

    Every decision, input, confidence value and action is written to an append-only log with cryptographic chaining. Any automated action can be reverted from the case timeline, restoring prior state and recording who reverted it and why.

Not ready to run it yourself?

Guardian MDR operates the same playbooks on your behalf with a three-minute response SLA, 24/7/365. You keep the tier configuration and the veto; our analysts carry the pager. Many customers start there and take operations in-house at the twelve-month mark.

Author your own

Playbooks are code, reviewed like code

Author in the visual editor or directly in the declarative language — they are the same artefact, so a diagram change produces a readable diff. Store playbooks in your own repository, gate them behind pull request review, and promote them through environments with the Guardian CLI.

playbook "contain-credential-theft" {
  trigger  detection.class == "credential_access"
  autonomy tier_for(asset.group, schedule.now)

  enrich {
    identity  = identity.risk(subject)
    asset     = inventory.lookup(host)
    intel     = intel.reputation(indicators)
  }

  guard {
    max_assets      = 25
    protected_lists = ["tier0", "ot-safety"]
  }

  decide confidence >= 0.88 and asset.tier > 0 {
    parallel {
      endpoint.isolate(host, keep_agent: true)
      identity.revoke_sessions(subject)
      identity.disable(subject, reversible: true)
      secrets.rotate(dependents_of(subject))
    }
    verify endpoint.rescan(host)
    close  with_evidence("full")
  } else {
    approval.request(
      channel: "slack://#soc-oncall",
      timeout: "15m",
      on_timeout: "execute_reversible"
    )
  }
}

Typed action library

1,200 actions across endpoint, identity, network, cloud, email, data and ticketing. Each declares its parameters, side effects, reversibility, required permission and expected latency — so the editor can validate a playbook before it ever runs.

Test fixtures included

Ship with recorded incident fixtures so a playbook has real unit tests. The CI action fails a pull request whose playbook changes behaviour on a fixture without an accompanying expectation update.

No lock-in on your logic

Playbooks export as plain text. Actions call documented REST endpoints you can invoke yourself. If you leave, your automation logic leaves with you in a form another engine can read.

See the API and SDK documentation

We ran everything in observe mode for six weeks and compared what Guardian would have done to what we actually did. It was right 340 times out of 347, and four of the seven misses were our own asset data being wrong. That report is what got automation approved by our risk committee.

Sanjay Kulkarni
SOC Manager, European retail bank

How teams roll this out

  1. Weeks 1–2

    Observe everything

    All playbooks run in Tier 0. You accumulate a decision record you can audit against your own analysts' choices.

  2. Weeks 3–6

    Assist on the noisy classes

    Move phishing, commodity malware and cloud misconfiguration to Tier 1. Triage time collapses first because that is where the wasted hours live.

  3. Weeks 7–10

    Supervise containment

    Tier 2 on endpoint and identity classes with approval gates on anything wide-scope. This is where MTTR falls off a cliff.

  4. Week 11 onward

    Autonomous where it earns it

    Promote to Tier 3 only the classes whose observed accuracy justifies it, reviewed quarterly against the reversal rate.

Questions

What SOC leaders ask us first

What happens if automation contains something it should not?

Every autonomous action is reversible and appears at the top of the case timeline with a rollback control. Isolation is lifted in under two seconds, a disabled account is restored with its prior group memberships, and rotated secrets are recoverable from the previous version in your secret manager. Across the customer base the reversal rate is 0.4%, and the most common cause is stale asset ownership data rather than a wrong verdict.

Do we have to replace our existing SOAR?

No. Sentinel exposes every detection, case and action over REST and webhooks, so an existing Splunk SOAR, XSOAR or Tines deployment can continue to own cross-domain orchestration while Sentinel handles security-domain containment where latency matters. Most customers gradually retire the playbooks that exist purely to gather context, because that context is already resolved.

How do you prevent an automated response from causing an outage?

Four mechanisms: scope ceilings that cap affected assets and rate, protected asset lists that exclude designated systems entirely, dependency analysis that names the services affected before an action runs, and change-freeze awareness that downgrades autonomy during declared windows. Anything that would breach a limit escalates for approval rather than proceeding.

Can we see why the AI Analyst reached a verdict?

Yes, and it is a hard requirement of the product. Every verdict shows the artefacts consulted, the detections correlated, the reasoning narrative, the confidence value and the specific factors that raised or lowered it. Any conclusion you disagree with can be corrected, and the correction feeds back into tuning for your tenant only — your data never trains a shared model.

Does automation work when our connection to Guardian is down?

Endpoint-local response — process termination, file rollback, host isolation — executes on the agent from a cached policy with no cloud dependency. Actions that require a third-party API queue and replay on reconnection with their original timestamps preserved. The platform itself carries a 99.99% availability commitment across three regional failover domains.

How long before we see a change in MTTR?

Triage time drops within the first fortnight simply from automated enrichment and case writing, typically 60–70%. Containment time changes the day you enable Tier 2 on endpoint and identity classes. The full 93% median reduction is measured at 90 days, once the noisy detection classes are promoted and the recovery workflows are wired into your ticketing.

Run it in observe mode and check our work

Six weeks of shadow execution against your live alerts produces a decision record you can compare to what your analysts actually did. It is the fastest way to find out whether autonomous response belongs in your SOC.