Detectiondetectionengineering

Detection Engineering

Telemetry becomes a detection hypothesis, an alert, an investigation and a response path; a noisy alert with no owner is not a control.

▶ Run the labFollow the failure

Frame the problem

Security starts with a concrete asset, attacker capability and trust crossing.

Asset
Time-to-detect and the evidence required to contain compromise.
Attacker & capability
An adversary using valid-looking actions after prevention fails.
Trust boundary
Security-relevant event → actionable responder signal
AssetThreatAttack SurfaceTrust BoundaryVulnerabilityExploit PathImpactMitigationDefense in DepthResidual Risk

Why the system fails

The system logs volume without identity/resource context, alerts on vague anomalies, or has no tested response attached.

The important question is not “what is Detection Engineering?” but “which assumption let untrusted data or an over-scoped identity cross security-relevant event → actionable responder signal?” Trace the decision at the boundary, then constrain what can happen after the first control fails.

Design the control in layers

Start with the control closest to the interpretation or privilege boundary: Design telemetry alongside the control Then add a control that reduces blast radius and telemetry that proves the decision was enforced.

The resulting design is not labelled secure. Record the identified controls, the known failure paths, the remaining exposure, and the evidence you would need during an incident.

PreventDetectRecover
Design telemetry alongside the control · Write detections as testable hypotheses with owners · Connect each alert to a containment actionCoverage tests, alert precision and time-to-triageContain the affected identity or component, scope impact from audit evidence, and preserve a regression test.

Key points

  • Asset: Time-to-detect and the evidence required to contain compromise.
  • Boundary: Security-relevant event → actionable responder signal
  • Primary control: Design telemetry alongside the control
  • Detection signal: Coverage tests, alert precision and time-to-triage
  • Always ask what limits damage when the primary control fails.

Boundary control exercise

This lesson uses the shared boundary-control exercise.

Boundary control check
Untrusted input / identity
Trust boundary
Privileged asset
Prevention may fail silently.

Follow the attack

Safe conceptual simulation: capability → missing control → crossed boundary → asset impact.

  1. 1
    Attacker starts with: An adversary using valid-looking actions after prevention fails.
  2. 2
    The system logs volume without identity/resource context, alerts on vague anomalies, or has no tested response attached.
  3. 3
    The weak or missing boundary control is crossed: Security-relevant event → actionable responder signal
  4. 4
    Impact: Attacker dwell time grows and incident scope remains unknown.
Blast radius
  • Attacker dwell time grows and incident scope remains unknown.

Defend, detect, recover

One prevention is a single point of security failure. Layer it and make failure observable.

Prevent
  • • Design telemetry alongside the control
  • • Write detections as testable hypotheses with owners
  • • Connect each alert to a containment action
Detect
  • • Coverage tests, alert precision and time-to-triage
Respond & recover
  • • Contain the affected identity or component.
  • • Scope access from audit evidence.
  • • Fix the boundary and add a regression test.
Residual risk
  • • Misconfiguration and new access paths can bypass the intended control.
  • • A privileged insider or compromised control plane may still reach the asset.

Misconceptions

Claim
“A single design telemetry alongside the control control makes this safe.”
Reality
One control changes risk; it does not erase it. Design prevention, detection, recovery, and blast-radius limits together.