AI Securityagenttrustboundaries

Agent Trust Boundaries

User, model, retrieval, memory, tools and external content have different trust and privilege; mark every flow explicitly.

▶ Run the labFollow the failure

Frame the problem

Security starts with a concrete asset, attacker capability and trust crossing.

Asset
The distinction between data, instructions, identity and authority inside an agent system.
Attacker & capability
Any actor controlling one context source.
Trust boundary
User/RAG/memory/tool output → model context → tool request
AssetThreatAttack SurfaceTrust BoundaryVulnerabilityExploit PathImpactMitigationDefense in DepthResidual Risk

Why the system fails

All text is flattened into one prompt and the model cannot reliably distinguish trusted policy from untrusted content.

The important question is not “what is Agent Trust Boundaries?” but “which assumption let untrusted data or an over-scoped identity cross user/rag/memory/tool output → model context → tool request?” Trace the decision at the boundary, then constrain what can happen after the first control fails.

Design the control in layers

Start with the control closest to the interpretation or privilege boundary: Track provenance and separate instructions from data structurally Then add a control that reduces blast radius and telemetry that proves the decision was enforced.

The resulting design is not labelled secure. Record the identified controls, the known failure paths, the remaining exposure, and the evidence you would need during an incident.

PreventDetectRecover
Track provenance and separate instructions from data structurally · Enforce tool policy after model output · Minimize context and capabilitiesLog provenance and policy decisions for every tool callContain the affected identity or component, scope impact from audit evidence, and preserve a regression test.

Key points

  • Asset: The distinction between data, instructions, identity and authority inside an agent system.
  • Boundary: User/RAG/memory/tool output → model context → tool request
  • Primary control: Track provenance and separate instructions from data structurally
  • Detection signal: Log provenance and policy decisions for every tool call
  • Always ask what limits damage when the primary control fails.

Boundary control exercise

This lesson uses the shared boundary-control exercise.

Boundary control check
Untrusted input / identity
Trust boundary
Privileged asset
Prevention may fail silently.

Follow the attack

Safe conceptual simulation: capability → missing control → crossed boundary → asset impact.

  1. 1
    Attacker starts with: Any actor controlling one context source.
  2. 2
    All text is flattened into one prompt and the model cannot reliably distinguish trusted policy from untrusted content.
  3. 3
    The weak or missing boundary control is crossed: User/RAG/memory/tool output → model context → tool request
  4. 4
    Impact: Untrusted data steers privileged actions.
Blast radius
  • Untrusted data steers privileged actions.

Defend, detect, recover

One prevention is a single point of security failure. Layer it and make failure observable.

Prevent
  • • Track provenance and separate instructions from data structurally
  • • Enforce tool policy after model output
  • • Minimize context and capabilities
Detect
  • • Log provenance and policy decisions for every tool call
Respond & recover
  • • Contain the affected identity or component.
  • • Scope access from audit evidence.
  • • Fix the boundary and add a regression test.
Residual risk
  • • Misconfiguration and new access paths can bypass the intended control.
  • • A privileged insider or compromised control plane may still reach the asset.

Misconceptions

Claim
“A single track provenance and separate instructions from data structurally control makes this safe.”
Reality
One control changes risk; it does not erase it. Design prevention, detection, recovery, and blast-radius limits together.