Constrain. Reason. Assure.
Architecture Note · Integrity for Agentic Systems
What can be enforced deterministically, where that ends, and how the residual is measured.
Draft for public comment · v0.7 · bounded.uk
This note does not claim to make untrusted content safe. It makes a narrower, defensible claim: a large class of agentic behaviour can be constrained deterministically; the boundary where that ends can be stated precisely; and the residual beyond it can be confined to a single declared gate and measured rather than trusted. Three movements — constrain, reason, assure.
Prompt injection has one structural cause: trusted instructions and untrusted text share a single token stream, and no model reliably acts on the first while merely reading the second. The danger isn't leaked secrets; it's low-integrity data contaminating high-integrity decisions — a Biba-style integrity problem. A source is where data enters; a sink is where it becomes a consequential effect — a tool call, a memory write, an outbound message, re-entry into trusted context. Injected content is harmless until it reaches a sink. The architectural mistake is asking the model to police itself; enforcement has to live somewhere verifiable, outside it. That framing — agentic AI as an integrity problem rather than a filtering one — is put most sharply by Schneier and Raghavan (2025), whose OODA-loop analysis was the spur for the approach taken here: “integrity isn’t a feature you add; it’s an architecture you choose.”
CBCR adopts the dual-path construction of Willison's dual-LLM pattern (2023) and CaMeL (Debenedetti et al., 2025): the component that reads untrusted content holds no capabilities; the component that reaches sinks never sees raw untrusted text. The contribution here is three things layered on top — naming exactly where construction is airtight and where it ends; a parametric residual-risk model for what's left; and a non-amplification rule extending the discipline to multi-agent delegation.
Enforcement is confined to a deterministic, verifiable reference monitor — the MCP Policy Firewall — and kept out of the model entirely. It checks labels and capabilities, never the meaning of content; that content-blindness is deliberate and defines the limit of what this part can promise. "Deterministic" stands in for verifiable. A permitted transition requires both capability and sufficient integrity, and either failure blocks the route.
flowchart TB
U["Untrusted / triaged input<br/>user prompt · RAG · tool output · memory · external agents"]
TZC["Trust Zone Classifier<br/>assigns integrity label"]
Q["Quarantined model<br/>holds no capabilities<br/>may see raw untrusted text"]
T["Typed, schema-validated channel"]
EG["Endorsement Gate<br/>only operation that raises integrity"]
P["Privileged model<br/>holds capabilities<br/>never sees raw untrusted text"]
K["High-integrity sinks<br/>tool call · memory write · redistribution · trusted re-entry"]
FW["MCP Policy Firewall<br/>deterministic reference monitor<br/>always invoked · tamper-resistant · verifiable"]
U --> TZC --> Q --> T --> EG --> P --> K
FW -. mediates .-> TZC
FW -. mediates .-> EG
FW -. mediates .-> P
FW -. mediates .-> K
classDef gate fill:#fbe4e2,stroke:#b23b34,stroke-width:2px,color:#1a1a1a;
classDef det fill:#e2ecfb,stroke:#3a66b0,color:#1a1a1a;
classDef cont fill:#efece6,stroke:#8a8275,color:#1a1a1a,stroke-dasharray:4 3;
class TZC,EG gate;
class FW,T,K det;
class U,Q,P cont;
Integrity is a lattice — untrusted ⊑ triaged ⊑ admissible ⊑ trusted — and the only operation that may raise an object's integrity is a declared, audited endorsement.
flowchart BT
UNT["untrusted"]
TRI["triaged"]
ADM["admissible"]
TRU["trusted"]
UNT --> TRI --> ADM --> TRU
UNT -. "endorsement<br/>(only upward operation —<br/>declared, constrained, audited)" .-> ADM
classDef low fill:#fbe4e2,stroke:#b23b34,color:#1a1a1a;
classDef mid fill:#fbf0d8,stroke:#c08a2e,color:#1a1a1a;
classDef high fill:#e1efe5,stroke:#3a8a55,color:#1a1a1a;
class UNT low;
class TRI mid;
class ADM,TRU high;
Two large classes can be constrained by construction, with no reliance on the model's judgement. Control-flow conformance: if the plan is derived from the trusted instruction and pinned as a contract, the firewall blocks any deviation — injected content can't move the agent off its plan, because the route doesn't exist. Closed-value-space data: where a call's parameters come from a trusted source and the response is a closed, validated shape — an enum, a range, an allowlisted ID, a strict format — validation is deterministic and "anything crafted doesn't match" is literally true.
The guarantee fails at a precise line, where three conditions must all hold: parameters from a trusted source, a trusted endpoint, and a closed value-space. The decisive one — schema validation checks shape, not meaning. A field typed string validates whatever it contains, so {recipient_email: "string"} passes identically whether it holds the intended address or the attacker's. The crafted value didn't break the structure; it is the structure, with hostile content. Type is not trust. Beyond that line — free text, semantically-loaded fields, data-derived parameters, untrusted endpoints — lies the residual.
The residual is handled first by construction: the instance that processes untrusted content holds no capabilities and emits only typed channels; the privileged instance that reaches sinks never sees raw untrusted text. Undetected instruction content can't reach a sink because the path doesn't exist — detection becomes defence-in-depth over a smaller surface, not the primary claim. But construction doesn't escape the open problem; it relocates it. The moment a typed channel carries a free-text field the validator can't semantically check, that field is the propagation problem wearing a schema.
flowchart TB
IN["untrusted content"]
LLM["model — black-box transformer<br/>no static analysis can track<br/>the flow through its weights"]
OUT["derived output<br/>carries no automatic label"]
RULE["only sound rule available:<br/>output from a context that held untrusted<br/>content is itself untrusted until endorsed"]
LOAD["cost: large volumes pushed through endorsement —<br/>the surface we wanted to minimise"]
TENSION["central tension:<br/>soundness of propagation ⇄ endorsement load"]
IN --> LLM --> OUT --> RULE --> LOAD --> TENSION
classDef low fill:#fbe4e2,stroke:#b23b34,color:#1a1a1a;
classDef box fill:#efece6,stroke:#8a8275,color:#1a1a1a;
classDef unk fill:#ffffff,stroke:#999999,color:#1a1a1a,stroke-dasharray:4 3;
classDef rule fill:#e2ecfb,stroke:#3a66b0,color:#1a1a1a;
classDef load fill:#fbf0d8,stroke:#c08a2e,color:#1a1a1a;
classDef tension fill:#fbe4e2,stroke:#b23b34,stroke-width:2px,color:#1a1a1a;
class IN low; class LLM box; class OUT unk;
class RULE rule; class LOAD load; class TENSION tension;
The only sound rule is conservative — any output from a context that held untrusted content is untrusted until endorsed — and that is what drives endorsement load up. Soundness ⇄ load is the central tension. Anyone who claims to have closed it should be asked for their propagation proof.
Whatever can't be made deterministic is confined to one place: the endorsement gate — the only operation that may raise integrity, and the maximally adversarial surface, since it reads attacker-controlled content to decide whether to trust it. There is no sound verification here, so three disciplines apply: fail closed (uncertainty refuses; endorsement may only reduce authority or pass); minimise the semantic surface (push every check you can into deterministic predicates, which returns them to Part I); and measure, don't trust (estimate the error rate by red-teaming and audit).
Under correct contracts and a sound enforcement layer, no untrusted source may influence a high-integrity sink except through a declared endorsement gate. This is noninterference modulo endorsement — a design property the architecture is built to satisfy, not a theorem proven over a formal semantics.
The residual is the unsoundness of the two probabilistic gates — the classifier that labels on the way in, and the endorsement gate that raises on the way out — weighted by the authority each can reach.
The deterministic core contributes nothing to this sum — by construction it has no probabilistic failure term. The lifecycle below shows where the gates fire and the two points a flow can halt; neither depends on catching malicious content.
sequenceDiagram
autonumber
participant SRC as Untrusted source
participant FW as MCP Policy Firewall<br/>(deterministic monitor)
participant TZC as Trust Zone Classifier
participant Q as Quarantined model<br/>(no capabilities)
participant EG as Endorsement Gate
participant P as Privileged model<br/>(holds capabilities)
participant K as Sink / Tool
SRC->>FW: inbound message
Note over FW: complete mediation —<br/>nothing enters except through here
rect rgb(251,228,226)
FW->>TZC: classify source
TZC-->>FW: label = untrusted / triaged
Note right of TZC: PROBABILISTIC gate —<br/>mislabel = label-soundness failure
end
FW->>Q: route untrusted content<br/>(Q can reach no sink)
Q-->>FW: typed, schema-validated output
Note over FW: conservative propagation:<br/>derived output stays untrusted until endorsed
rect rgb(251,228,226)
FW->>EG: request endorsement to admissible
Note right of EG: PROBABILISTIC gate —<br/>false-endorsement = residual-risk locus
alt endorsement granted
EG-->>FW: raise label to admissible
else endorsement refused
EG-->>FW: reject
FW-->>SRC: flow halted — no sink reached
end
end
FW->>P: deliver admissible typed data<br/>(P never sees raw untrusted text)
P-->>FW: proposed action — invoke tool with args
Note over FW: deterministic check before any sink
alt capability held AND minimum integrity met
FW->>K: invoke tool
K-->>FW: result
Note over FW: result re-enters as a NEW source<br/>and is classified again
else check fails
FW-->>P: denied (authority-reducing only)
end
For multi-agent systems, a monotonic non-amplification rule keeps trust from growing along a chain: each hop's capabilities and reachable sinks are a subset of the delegator's, checked by the firewall against a proof-carrying contract chain. It bounds authority, not information — what a compromised agent can do, not what it can say to its delegees.
flowchart LR
A["Agent A (origin)<br/>caps = {read, write, send, delete}<br/>reachable sinks = S_A"]
B["Agent B<br/>caps(B) ⊆ caps(A)<br/>sinks(B) ⊆ S_A"]
C["Agent C<br/>caps(C) ⊆ caps(B)<br/>sinks(C) ⊆ sinks(B)"]
A -- "delegate · proof-carrying<br/>firewall checks attenuation" --> B
B -- "delegate · proof-carrying<br/>firewall checks attenuation" --> C
classDef a fill:#e2ecfb,stroke:#3a66b0,color:#1a1a1a;
classDef b fill:#fbf0d8,stroke:#c08a2e,color:#1a1a1a;
classDef c fill:#fbe4e2,stroke:#b23b34,color:#1a1a1a;
class A a; class B b; class C c;
This matters increasingly as agents begin to discover one another at runtime — the Linux Foundation's DNS-AID (2026) publishes and verifies agents and MCP servers over DNS. Discovery and identity answer who an agent is and where to reach it; even verified, that is authenticity, not content-safety. CBCR governs the layer above — what a discovered agent may do, and how its content is kept from a sink. Every newly-discovered counterpart is an untrusted endpoint until contract and integrity say otherwise.
CBCR does not make untrusted content safe. It constrains what can be constrained deterministically, contains the model so it cannot reach a sink directly, and confines the irreducible residual to one declared, fail-closed, measured gate — bounding the authority of whatever crosses it. A smaller claim than "the firewall stops malicious payloads," and a stronger one.
The dual-path construction — capability-separated, untrusted-content-quarantined execution — from Willison's dual-LLM pattern (2023) and CaMeL (2025).
(i) The deterministic/residual carve-out — naming where construction is airtight and exactly where it ends (type is not trust). (ii) A parametric residual-risk model localising and measuring what is left, at two named gates. (iii) A monotonic non-amplification rule extending the discipline to multi-agent chains.
Conservative label propagation through the model itself. Stated, not solved.