Conformance Verification is the section every other section defers to. Each prerequisite and invariant states what an implementation must do and closes with the tests that demonstrate it does; §22 specifies what those demonstrations must establish, and by what method, so that conformance to ZTG is itself a checkable fact rather than an asserted one. This is not a peripheral appendix. The framework's central claim is that it produces checkable governance — governance an auditor, a counterparty, or a court can independently verify, not merely a system's account of its own good behavior — and that claim is only as strong as the verification §22 defines. Until conformance can be demonstrated, every "Conformance Criteria" list in the preceding chapters is a statement of intent.
Operational Questions #
§22 is not mapped to one of the seven operational questions; it is the section that makes all seven verifiable. Each question is answered by some prerequisite or invariant, and §22 specifies how an implementation demonstrates that the answer holds: how it shows admissibility is actually established before a grant, that fail-closed semantics actually fire, that replay actually reproduces, that the effect surface is actually closed. The section's subject is therefore the verifiability of the whole, and its discipline is that a guarantee which cannot be demonstrated is, for governance purposes, not yet a guarantee.
Normative #
An implementation is ZTG-conformant to the extent that it satisfies the structural prerequisites (ZTG-0a–0e) and the system invariants (ZTG-1–5), and can demonstrate that it does so by the methods specified in this section. Satisfaction and demonstration are distinct requirements: an implementation that meets a requirement but cannot show it meets the requirement has not established conformance to that requirement, because the property the framework sells is verifiability, and an unverifiable guarantee is outside what the framework certifies.
Conformance is not self-certifying by assertion. A claim of conformance MUST be backed by the evidence and procedures this section requires, such that a party who does not trust the implementer can reach the same conclusion from the same evidence. The substrate that makes this possible is the same substrate the prerequisites require: because decisions are recorded (ZTG-0a), reproducible (ZTG-0b), and attributable (ZTG-0d), a third party can re-derive conformance findings rather than accept them.
Property Classes and Their Verification Methods #
The requirements across ZTG are not all verified the same way, and a single uniform method — "run a test suite" — is insufficient for several of them. §22 distinguishes the classes of property the framework asserts and specifies the method each admits.
Reachability and structural properties. Some requirements assert the absence of a path: no agent-to-execution route bypasses the boundary (ZTG-1); no reachable effect channel is unregistered (ZTG-3); no ambient effect authority exists (ZTG-3). These cannot be established by testing, because testing exercises paths that exist and cannot demonstrate that no other path does. They MUST be verified by architecture-level reachability analysis: a demonstration over the implementation's structure that the only reachable paths to the world traverse the mediated, registered surface. An implementation MUST be able to present its effect-reachability argument, not merely a suite of passing effect tests.
Reproduction properties. Some requirements assert that a result regenerates: replay reproduces a verdict (ZTG-0b); banded assessment recomputes identically (ZTG-5); an admissibility determination is independently reproducible (§3). These are verified by re-execution over a recorded corpus and comparison: regenerate the determination from its substrate under the pinned procedure and confirm identity with the recorded result.
Negative-space and absence properties. Some requirements assert that missing things are detected: an expected record absent on an action path is a violation (ZTG-0a); an effect lacking committed evidence is an orphan to be caught (ZTG-4). These are verified by injection — deliberately producing the absent condition and confirming the implementation raises it as a violation rather than passing it as ordinary missing data. For the ZTG-4 no-orphan direction specifically, the demonstration is that completion records are emitted only after dispatch on the isolated coupling path, so the absence of a corresponding effect for a completion record cannot arise silently (per §13's authorizing-vs-completion division).
Fault-injection properties. Some requirements assert a response to loss of guarantee: Stasis fires on each trigger (ZTG-2); a partitioned gate fails closed (ZTG-0e); the write-ahead crash window yields a recorded indeterminate effect, never a silent gap (ZTG-4). These are verified by injecting the fault and confirming the fail-closed response — including that recovery from Stasis requires ratified authority and, for tamper-family triggers, independent re-verification (ZTG-2 Exit Path Integrity).
Integrity properties. Some requirements assert tamper-evidence: governance records cannot be silently altered (ZTG-0a); the coupling path cannot be severed by the reasoning system (ZTG-4). These are verified by adversarial attempt — altering, deleting, or inserting records, and attempting to dispatch an effect while suppressing its evidence — and confirming detection.
An implementation's conformance evidence MUST address each requirement by the method its property class admits. A reachability property demonstrated only by passing tests, or an integrity property demonstrated only by analysis, is not adequately demonstrated.
Closure Verification #
The hardest structural demonstration is ZTG-3 closure: that no reachable effect channel is unregistered. §22 specifies it as a reachability obligation. The implementation MUST exhibit its complete set of effect channels and an argument that every reachable path from the governed system to an externally observable effect terminates in a registered surface — equivalently, that the capacity to produce an unregistered effect does not exist as an ambient capability (ZTG-3). The argument is over the architecture, and its completeness is the conformance burden: ZTG-3 reframes "have we listed every effect" as "is there any reachable path outside a registered surface," and §22 is where that reachability question is discharged for a given implementation. Discovery, by analysis or at runtime, of a reachable unregistered channel is a closure breach and a Stasis trigger (ZTG-3), not a finding to be quietly registered away.
The Convergence Test #
The convergence test is the method by which an admissibility determination (§3) — and the harm classification within it (ZTG-5) — is shown to be checkable rather than asserted. The test requires that an independent assessment of a request, performed over the recorded substrate and not sharing state with the original evaluation, converge on the same determination. Convergence is the conformance evidence for §3's claim that admissibility is reproducible; divergence between independent assessments of the same recorded request is a conformance failure, surfaced for investigation.
The independent assessment MUST draw only on the recorded substrate — recorded inputs, pinned policy and engine, pinned snapshot — and MUST NOT consult the live system or re-run the reasoning model, for the same reasons replay does not (ZTG-0b): an assessment that re-derived its inputs from the live system would be checking the system against itself. The precise criteria for convergence — exact-match on verdict and harm class, and the tolerance, if any, permitted on banded composite assessment — are specified below as an open calibration item; the requirement that independent assessment converge, and that non-convergence be a conformance failure, is settled.
Conformance of Non-Reference Implementations #
The specification is open: Constable is the reference implementation, and any implementation satisfying the prerequisites and invariants is ZTG-conformant (Scope). §22 is what makes that claim operational for a non-reference implementation. A non-reference implementation demonstrates conformance by exhibiting, for each requirement, the evidence its property class requires — its reachability arguments, its reproduction corpus and harness, its injection and adversarial results — over its own architecture. Conformance is to the specified invariants and prerequisites, not to Constable's mechanisms: an implementation that achieves complete mediation, closure, evidence coupling, and the rest by different means than Constable is conformant if its evidence establishes the properties, and Constable's particular choices (OPA/Rego, the Monotonic Logger, HumanSeal) are not themselves conformance requirements.
Scope and Partial Conformance #
An implementation's conformance claim MUST state the scope over which it holds — which effect surfaces, which action classes, which deployment. A demonstration that covers a subset of an implementation's effectful actions establishes conformance for that subset only; the unscoped remainder is not conformant by extension. Consistent with the framework's posture elsewhere, §22 does not soften a requirement to fit an implementation not yet able to demonstrate it: an implementation that cannot yet exhibit closure over its full effect surface is partially conformant over the surface it can, not fully conformant over all of it. Silent scope limitation — claiming conformance while bounding coverage without saying so — is itself a conformance defect.
Conformance Criteria #
A conforming verification regime can: classify each ZTG requirement by property class and apply the admitting method; present effect-reachability arguments for the structural properties (no-bypass, closure, no-ambient-authority); regenerate recorded determinations under pinned procedure and confirm reproduction; inject negative-space and fault conditions and confirm violation-handling and fail-closed response; attempt tamper and confirm detection; demonstrate independent convergence on recorded admissibility determinations; and state the scope over which the conformance claim holds, without silent limitation.
Further Considerations #
The assurance-case grounding. Safety-critical and high-assurance engineering long ago stopped accepting "we tested it" as a sufficient account of why a system can be trusted, and developed the assurance case: an explicit, auditable argument that a system meets its claims, structured so the claims decompose into sub-claims and bottom out in evidence a reviewer can examine. Independent-evaluation regimes — the kind that certify a cryptographic module or an avionics component — institutionalize the same idea: the developer does not certify their own product by assertion; an independent party reaches the conclusion from the evidence. §22 asks an implementation to assemble exactly this: a structured argument that each ZTG requirement holds, decomposed by property class, bottoming out in reachability analyses, reproduction corpora, and injection results a party who distrusts the implementer can re-examine. The tradition's central lesson is the one §22 encodes — that testing demonstrates the presence of behavior, not its absence, so the absence claims (no bypass, no unregistered channel, no silent gap) require argument and analysis, not test counts.
Verifiability is the product. It is worth stating plainly why this section carries the weight it does. A system can satisfy every ZTG requirement and still deliver little if it cannot show that it does, because the institutional value the framework offers — to a regulator, an underwriter, a court — is not that the system behaves well but that its behavior is demonstrable. Governance that must be taken on the operator's word is the retrospective, advisory posture the introduction rejects (§1.2), dressed in technical language. §22 is where the framework's difference from that posture is cashed: the guarantees are constructed so that a third party can confirm them, and the confirmation procedure is itself specified rather than left to the implementer's discretion.
Self-certification and its limits. An implementation can run its own conformance regime, and Constable does. But self-administered evidence is only as trustworthy as the substrate it rests on, which is why the framework invests so heavily in records that are tamper-evident (ZTG-0a) and determinations that are independently reproducible (ZTG-0b, §3): these are what let a self-administered claim be re-checked by an outside party rather than merely believed. The framework's posture is referee-shaped: it specifies the conditions under which conformance can be independently confirmed, so that self-certification and third-party evaluation rest on the same evidence and reach the same finding.
Reachability completeness is irreducibly architectural. The recurring hard problem, named in ZTG-1 and ZTG-3 and concentrated here, is establishing an absence-of-path claim over a real system. §22 does not pretend this is mechanical. The completeness of a reachability argument depends on the implementation's architecture being analyzable — which is itself a design property an implementation must choose, and a reason the framework prefers architectures whose effect surface is small, explicit, and structurally closed rather than large and conventionally monitored. An implementation that makes its effect reachability hard to analyze has made its conformance hard to establish, and the framework treats that as the implementation's burden, not a gap in §22.
How We Do It #
Constable's conformance regime assembles the per-chapter test obligations into a single suite organized by property class, and pairs the testable obligations with the analyses the structural properties require.
Per-property suites. Constable runs the conformance tests each chapter specifies — ZTG-0a coverage/tamper/gap/metric tests, ZTG-0b reproduction/drift/isolation tests, ZTG-0c time tests, ZTG-0d identity tests, ZTG-0e atomicity/partition tests, ZTG-1 bypass/interception tests, ZTG-2 trigger/exit tests, ZTG-3 reachability/breach tests, ZTG-4 ordering/orphan/crash tests, ZTG-5 classification/banding/promotion tests, and §3 admissibility/verdict/provenance tests — and organizes them by the property class each falls into rather than by chapter alone.
Reachability analysis. For the structural properties, Constable maintains an effect-reachability argument over its architecture: the agent runtime holds no ambient effect capability, every outbound effect routes through a registered surface adaptor, and the gate is the only path from proposal to dispatch. This argument, not a test pass, is what discharges the no-bypass (ZTG-1) and closure (ZTG-3) obligations; a reachable effect path found without a registered surface raises an integrity violation and triggers Stasis.
Reproduction and convergence harness. Constable's replay harness (ZTG-0b) regenerates recorded determinations under pinned bundle and engine; the convergence regime runs independent assessment over the recorded substrate and compares, surfacing divergence as a conformance failure. Both operate without reaching the effect surface.
Injection and adversarial harness. Constable injects negative-space, fault, and tamper conditions — missing records, write-ahead crashes, altered log entries, partition between gates, suppressed evidence — and confirms the specified violation-handling and fail-closed responses.
Certification. Forthcoming Constable certification packages this regime so that a deployment's conformance, and the scope over which it holds, can be presented as auditable evidence rather than asserted. The artifacts are constructed for re-examination by an outside party, consistent with the referee posture above.