The Input Sanitization Boundary is the input-side composition requirement (§4) detailed in full. ZTG-1 states the limit it answers: the boundary cannot be more robust than the inputs it evaluates. A deterministic gate evaluating malformed, ambiguous, or adversarially-shaped input can produce a correct verdict on a misread of what was actually proposed. §18 is the requirement that inputs are brought to a canonical, validated form before evaluation, so that the gate evaluates well-formed inputs whose meaning is unambiguous, and so that the form a later replay reconstructs is the form the gate actually saw.
Operational Questions #
§18 is not one of the seven operational questions; it is a precondition for two of them. It serves execution-boundary enforcement (with ZTG-1 and ZTG-3) by ensuring the boundary mediates inputs whose meaning is determinate rather than inputs an adversary can shape to be read two ways. And it serves replayability (ZTG-0b): the canonical, post-sanitization input is the input the gate evaluated and the input replay reconstructs, which is why §6 records the post-normalization input or pins the normalization that produced it. An un-sanitized input path is both an enforcement gap and a replay gap.
Normative #
Every input that reaches governance evaluation MUST first pass the Input Sanitization Boundary, and the gate MUST accept only sanitized inputs. There is no input path into evaluation that bypasses sanitization. Sanitization brings an input to a canonical validated form, rejects input that cannot be brought to such a form, and records what it did.
Normalization to Canonical Form #
Inputs MUST be normalized to a canonical representation before evaluation, such that semantically-equivalent inputs normalize to the same form and evaluate identically. The purpose is to remove the representational ambiguity adversarial inputs exploit: encoding tricks, equivalent-but-different spellings of a target, hidden or duplicated fields, and the many ways one effective request can be made to look like another. A policy evaluated against a non-canonical input is evaluating against one reading of an input that has several, and the reading the gate takes need not be the reading that produces the effect. Canonicalization collapses those readings to one before the gate decides.
Validation and Rejection #
Input that cannot be normalized to a valid canonical form MUST be refused, not evaluated on a best-effort basis. Sanitization is fail-closed: an input the boundary cannot bring to a well-formed, validated representation is rejected as malformed, and the rejection is recorded. The architecture does not guess at the intended meaning of a malformed input and proceed on the guess, because a guess is exactly the ambiguity an adversary supplies the input to create.
Model Output Is Untrusted Input #
The reasoning system's output is input, and it is the central case. Under §4 the reasoning system composes as a proposer outside the trust boundary, so its outputs — tool-call proposals, action parameters, target specifications — are untrusted input and MUST pass sanitization before they enter an authorization request (§3). This is where the structural attack surface of prompt injection and manipulated model output is addressed: §18 does not make the model trustworthy, and it does not certify that the model's proposal reflects the user's intent; it ensures the proposal reaches the gate in canonical, validated, bounded form, so that whatever the model was induced to emit is evaluated as the well-formed action it actually is, against policy, at its true harm class. The model's output earns no exemption from sanitization by virtue of originating inside the deployment.
Sanitization Is Not Authorization #
§18 validates form; it does not decide admissibility. A perfectly sanitized input can be, and often is, refused by policy (§3): canonical form is a precondition for the gate to decide correctly, not a determination that the input should be granted. The two MUST NOT be conflated. Sanitization that began making policy decisions would be a second, undocumented authorization surface; admissibility belongs to the authorization model, and §18 delivers it inputs it can decide on, nothing more.
Determinism and Provenance #
The normalization transform MUST be deterministic, and either its output (the
post-normalization input) MUST be recorded as the replay input or the transform version
MUST be pinned alongside policy and engine (ZTG-0b, §6). A nondeterministic or
silently-drifting normalizer would make verdicts non-reproducible for reasons unrelated to
any governance decision. Each sanitization MUST emit input-normalization evidence
(INPUT_NORMALIZED, ZTG-0a) carrying the provenance of the sanitized input: what raw
input it derived from and what normalization was applied, sufficient to reconstruct the
gate's input under replay.
Scope #
Sanitization applies to every input class that reaches evaluation: end-user input, reasoning-system output, promoted memory content (ZTG-1), and external data drawn into the decision. The boundary-scope subtleties ZTG-1 names — reads whose access pattern signals a target, mid-generation tool calls, memory writes consulted elsewhere — are input surfaces as much as effect surfaces, and an input crossing into evaluation through any of them is subject to §18.
Conformance Criteria #
A conforming implementation can: demonstrate that no input reaches evaluation without
passing sanitization; demonstrate canonicalization, such that semantically-equivalent
inputs normalize identically; refuse malformed and un-normalizable input fail-closed and
record the refusal; sanitize reasoning-system output as untrusted input before it enters a
request; demonstrate that sanitization makes no admissibility decision; demonstrate that
its normalizer is deterministic and that the replay input is the canonical form or the
normalizer is pinned; and emit INPUT_NORMALIZED provenance sufficient for replay.
Further Considerations #
The language-theoretic-security grounding. A line of security research traced a large class of vulnerabilities to a single habit: processing input before, or while, deciding whether it is well-formed — the "shotgun parser" that interleaves recognition and action, so that malformed or ambiguous input is partway processed before anyone has established what it is. The discipline that answers it treats input as a formal language, places a recognizer at the boundary that accepts only well-formed input and rejects everything else before any processing, and refuses the equivalence-ambiguities that let one input masquerade as another. §18 is this discipline at the governance boundary. Canonicalization is the recognizer bringing input to one well-formed representation; fail-closed rejection is the refusal to process what the recognizer cannot accept; the insistence that sanitization precede evaluation is the refusal of the shotgun parser. The tradition's core finding is the one §18 depends on: robustness against adversarial input is a property of validating form fully at the boundary, not of handling malformed input gracefully downstream.
What sanitization can and cannot do. §18 removes a structural attack surface; it does not resolve a semantic one, and the chapter is precise about the line. Canonicalization and validation ensure the gate evaluates the well-formed action an input actually encodes — they defeat the ambiguity and malformation vectors. They do not, and cannot, certify that a well-formed model proposal reflects what a user intended or that a manipulated model has not been induced to make a well-formed but unwanted request. That residual is addressed elsewhere in the architecture: by policy refusing the action (§3), by harm-class gating routing it to human authority (ZTG-5), and by alignment work in the Envelope. §18's contribution is bounded and stated as bounded — it makes the input determinate; it does not make the proposer well-intentioned.
Why the model's output is where the work is. Among input classes, reasoning-system output is the highest-volume and the most adversarially-influenced, because it is shaped by whatever entered the model's context, including hostile content. Treating it as untrusted input rather than as a privileged internal signal is the single most consequential stance in §18. An architecture that sanitized external user input but trusted model output would have sanitized the smaller surface and exempted the larger one.
Relationship to replay. The canonical form §18 produces is the anchor §6 relies on: by recording the post-normalization input (or pinning the normalizer), the architecture makes the gate's input reconstructable, so a replayed decision evaluates the input the gate saw rather than a re-derivation of it. §18 and ZTG-0b meet exactly here, and the #8 reconciliation in §6 is the other half of this requirement.
How We Do It #
Constable implements §18 as Airlock, the input-sanitization component every input traverses before the execution gate.
Canonicalization and validation. Airlock normalizes inputs to a canonical form and validates them against the expected structure for their input class, rejecting malformed or un-normalizable input fail-closed. The gate accepts only Airlock-processed inputs; there is no path by which raw input reaches policy evaluation.
Model-output sanitization. Tool-call proposals and action parameters produced by the agent runtime pass through Airlock as untrusted input before the gate evaluates them. Constable derives the harm-determining features used for routing and classification (§14, §12) from the Airlock-validated parameters, never from a model-asserted tag — the validated form is what the gate and the ZTG-5 routing function read.
Determinism, provenance, and replay. Airlock's normalization is deterministic; Constable
records the post-normalization input as the replay input and emits INPUT_NORMALIZED
(ZTG-0a) with the provenance of each sanitized input, so a ZTG-0b replay reconstructs the
gate's input exactly. Normalization that is version-relevant is pinned with the governance
bundle.
Conformance tests. Constable's testing for §18 includes: bypass tests confirming no input reaches the gate without Airlock; canonicalization tests confirming equivalent inputs normalize identically and encoding/ambiguity tricks collapse to one form; rejection tests confirming malformed input is refused fail-closed and recorded; model-output tests confirming agent output is sanitized as untrusted input and that routing features derive from validated parameters; and replay tests confirming the recorded canonical input reconstructs the gate's input. The protocol is documented in the conformance verification specification referenced in §22.