Agent Tool Capability Security
A tool schema is a capability interface: make it narrow, bind it to a principal, enforce policy outside the model and record the result.
Frame the problem
Security starts with a concrete asset, attacker capability and trust crossing.
Why the system fails
A generic shell/HTTP/database tool or admin token lets the model select arbitrary authority and target.
The important question is not “what is Agent Tool Capability Security?” but “which assumption let untrusted data or an over-scoped identity cross model-generated arguments → deterministic privileged operation?” Trace the decision at the boundary, then constrain what can happen after the first control fails.
Design the control in layers
Start with the control closest to the interpretation or privilege boundary: Expose task-specific tools with typed narrow arguments Then add a control that reduces blast radius and telemetry that proves the decision was enforced.
The resulting design is not labelled secure. Record the identified controls, the known failure paths, the remaining exposure, and the evidence you would need during an incident.
| Prevent | Detect | Recover |
|---|---|---|
| Expose task-specific tools with typed narrow arguments · Authorize the real principal, action and resource outside the model · Separate read/write/destructive capabilities and require approval | Audit proposed, denied, approved and executed actions | Contain the affected identity or component, scope impact from audit evidence, and preserve a regression test. |
Key points
- Asset: Money, files, accounts, communications and infrastructure reachable through tools.
- Boundary: Model-generated arguments → deterministic privileged operation
- Primary control: Expose task-specific tools with typed narrow arguments
- Detection signal: Audit proposed, denied, approved and executed actions
- Always ask what limits damage when the primary control fails.
Boundary control exercise
This lesson uses the shared boundary-control exercise.
Follow the attack
Safe conceptual simulation: capability → missing control → crossed boundary → asset impact.
- 1Attacker starts with: Any context source able to influence the model’s tool proposal.
- 2A generic shell/HTTP/database tool or admin token lets the model select arbitrary authority and target.
- 3The weak or missing boundary control is crossed: Model-generated arguments → deterministic privileged operation
- 4Impact: Destructive or cross-user actions at machine speed.
- Destructive or cross-user actions at machine speed.
Defend, detect, recover
One prevention is a single point of security failure. Layer it and make failure observable.
- • Expose task-specific tools with typed narrow arguments
- • Authorize the real principal, action and resource outside the model
- • Separate read/write/destructive capabilities and require approval
- • Audit proposed, denied, approved and executed actions
- • Contain the affected identity or component.
- • Scope access from audit evidence.
- • Fix the boundary and add a regression test.
- • Misconfiguration and new access paths can bypass the intended control.
- • A privileged insider or compromised control plane may still reach the asset.