AI Securityragandagentmemorysecurity

RAG and Agent Memory Security

Retrieval and memory add durable, searchable copies of data where poisoning, tenant-filter mistakes, retention and sensitive recall become security boundaries.

▶ Run the labFollow the failure

Frame the problem

Security starts with a concrete asset, attacker capability and trust crossing.

Asset
Private documents, tenant isolation and the integrity of future agent context.
Attacker & capability
A document author, another tenant or a user able to write persistent memory.
Trust boundary
Document/memory write → retrieval → another decision or user
AssetThreatAttack SurfaceTrust BoundaryVulnerabilityExploit PathImpactMitigationDefense in DepthResidual Risk

Why the system fails

Metadata filtering is optional, malicious content persists, or retention and read access are unspecified.

The important question is not “what is RAG and Agent Memory Security?” but “which assumption let untrusted data or an over-scoped identity cross document/memory write → retrieval → another decision or user?” Trace the decision at the boundary, then constrain what can happen after the first control fails.

Design the control in layers

Start with the control closest to the interpretation or privilege boundary: Enforce tenant filters before retrieval Then add a control that reduces blast radius and telemetry that proves the decision was enforced.

The resulting design is not labelled secure. Record the identified controls, the known failure paths, the remaining exposure, and the evidence you would need during an incident.

PreventDetectRecover
Enforce tenant filters before retrieval · Separate write trust from read/use trust · Set retention, provenance and deletion controlsCross-tenant retrieval tests and provenance anomaliesContain the affected identity or component, scope impact from audit evidence, and preserve a regression test.

Key points

  • Asset: Private documents, tenant isolation and the integrity of future agent context.
  • Boundary: Document/memory write → retrieval → another decision or user
  • Primary control: Enforce tenant filters before retrieval
  • Detection signal: Cross-tenant retrieval tests and provenance anomalies
  • Always ask what limits damage when the primary control fails.

Boundary control exercise

This lesson uses the shared boundary-control exercise.

Boundary control check
Untrusted input / identity
Trust boundary
Privileged asset
Prevention may fail silently.

Follow the attack

Safe conceptual simulation: capability → missing control → crossed boundary → asset impact.

  1. 1
    Attacker starts with: A document author, another tenant or a user able to write persistent memory.
  2. 2
    Metadata filtering is optional, malicious content persists, or retention and read access are unspecified.
  3. 3
    The weak or missing boundary control is crossed: Document/memory write → retrieval → another decision or user
  4. 4
    Impact: Cross-tenant disclosure, persistent prompt injection or sensitive-data recall.
Blast radius
  • Cross-tenant disclosure, persistent prompt injection or sensitive-data recall.

Defend, detect, recover

One prevention is a single point of security failure. Layer it and make failure observable.

Prevent
  • • Enforce tenant filters before retrieval
  • • Separate write trust from read/use trust
  • • Set retention, provenance and deletion controls
Detect
  • • Cross-tenant retrieval tests and provenance anomalies
Respond & recover
  • • Contain the affected identity or component.
  • • Scope access from audit evidence.
  • • Fix the boundary and add a regression test.
Residual risk
  • • Misconfiguration and new access paths can bypass the intended control.
  • • A privileged insider or compromised control plane may still reach the asset.

Misconceptions

Claim
“A single enforce tenant filters before retrieval control makes this safe.”
Reality
One control changes risk; it does not erase it. Design prevention, detection, recovery, and blast-radius limits together.