Experiment · Paper 2 · Delegation-Grade Agents v1.1
EXP-006 — Confirmatory E1+E2
Preregistered 540-run replication of outcomes-only insufficiency and behavior-contract dominance across the four canonical OPBR domains. The Paper 1 effects survive tightened conditions, frozen seeds, and an explicit blind on the perturbation set.
Preregistered hypotheses
E1: Under canonical OPBR scenarios held out from Paper 1 training, output-only evaluation will fail to separate regression from baseline runs at any operating point.
E2: Behavior contracts will achieve F1 ≥ 0.90 on the same scenarios, with ΔF1 ≥ 0.30 over the best non-contract baseline.
Method
540 runs across the four OPBR domains under the tightened Paper 2 protocol. Perturbation set is frozen prior to execution and recorded in PAPER2_PROTOCOL_FREEZE.json. Seeds are bound to scenarios and the agent does not have access to any of the strings in the registered failure-mode lexicon. Detector evaluation is paired McNemar against every baseline.
Result
E1 confirmed: output-only evaluation passes both regression and baseline at indistinguishable rates under the held-out perturbations. E2 confirmed: contracts cross the preregistered F1 floor and the ΔF1 floor against every baseline (Holm-corrected).
What this experiment does not establish
This confirmatory experiment establishes replication of outcomes-only insufficiency and contract dominance on the canonical OPBR distribution, with one model and one runtime. It does not establish generalization to runtimes other than Claude Code (Codex replication pending in EXP-007), and it does not test out-of-distribution failure modes outside the four OPBR domains.
Replication
abw run --experiment EXP-006
Requires Agent Behavior Workbench ≥ 0.4. Corpus auto-downloads from OPBR-Bench v0. Full protocol freeze at exp-006-protocol.json.
Artifacts
Every load-bearing claim on this page traces back to one of the artifacts below. They are the canonical citation targets — not the prose.
exp-006-protocol.json— frozen preregistration: hypothesis, falsifier, method, detectors, primary metric, preregistered gates, claim boundary.exp-006-results.json— machine-readable result block: primary metric, gate-pass record, mechanism interpretation.- Agent Behavior Workbench — open-source code and OPBR-Bench v0 corpus. Required to run
abw run --experiment EXP-006. - Paper 2 · Delegation-Grade Agents · v1.1 — Whitepaper anchored on this confirmatory experiment.
Cite as: Sanchez de la Sierra, A. (2026). EXP-006 — Confirmatory E1+E2. Raising Agents Lab. https://raisingagents.is/lab/exp/exp-006
Related
- Paper 2: Delegation-Grade Agents — Why Determinism Is Not Enough for Agent Reliability
- Prior: EXP-004 — Held-out repair (Paper 1 Study 5, the experiment this confirms)
- Pending: EXP-007 — Cross-runtime (Codex CLI)