Adrian's field letter · bi-weekly · public-safe
Behavior Watch.
A bi-weekly field letter for engineers and leaders trying to delegate significant work to agents. Each issue names one failure pattern, shows one public-safe artifact, and asks what would make the work safe to authorize: owner, evidence, threshold, rollback, receipt, review. One failure pattern. One artifact. One delegation boundary. Private team questions route through Zartis where appropriate.
Follow the ongoing thinking.
Free. Bi-weekly. No sponsors. One unsubscribe link in every issue. The publication uses public, synthetic, redacted, or explicitly approved artifacts only. Selected Free Experiments may appear here when the evidence is strong enough.
Free Experiments
Have a behavior problem worth testing?
Submit it to the lab. The intake agent will help turn the problem into a falsifiable, public-safe experiment idea. Selected ideas may become free lab experiments or Behavior Watch case files.
The case-file format
Most issues follow the same six-section structure. The same shape every time so a senior engineer can scan it in 90 seconds or read it in 12 minutes — by choice, not by accident.
/ TL;DR
Up to 250 words. The claim, the shape of the evidence, the verdict. Read this if you are skimming.
/ Case file
Full analysis. Length determined by the case, not a reading-time target. 2,500–6,000 words typical.
/ Artifacts
A list of files in the issue's artifact pack. Behavior contracts, trace pairs, evaluator reports, scripts. Downloadable.
/ Reproduction
A command-line sequence the reader can run against the artifact pack. With workbench installed, under a minute.
/ Sources
Minimum five per issue. At least 60% primary sources (papers, source code, official docs).
/ Delegation boundary
Short statement of what the agent may do, when it must ask, when it must escalate, and what evidence would change the answer.
Issue types
Six recurring formats. Every issue is one of these, but the question underneath is stable: what does this teach us about safe delegation?
- 01 · Paper-to-practiceA new paper applied to a concrete production case. What changes if you adopt this?
- 02 · Trace teardownOne trace pair. Same final output, two different paths. The signature format of the publication: plausible is not permission.
- 03 · Failure patternA class of failure observed across multiple deployments. Named, characterised, with detection guidance.
- 04 · Build artifactOne concrete workbench component or contract specification, shipped with full reproduction.
- 05 · Tool / vendor watchEval / observability vendor change, audited for what it does and does not detect.
- 06 · Claim-safety noteA claim from a paper or product page audited against its evidence. What survives, what doesn't.
Issue pipeline
This is a publication pipeline, not a launch calendar. Published issues stay visible, corrections stay attached to the record, and candidates ship only when the public artifact is strong enough. Team-specific material is routed through Zartis instead of being turned into public content.
-
BW-2026-000
Type 00 · orientationcandidateorientationHow this newsletter works.
The orientation issue. The case-file format demonstrated on itself: TL;DR, case file, artifacts, reproduction, sources, claim boundary. Reader gets the schema and the publication boundary before the next artifact lands.
-
BW-2026-001
Type 02 · teardowncandidatetrace teardownThe €50K refund that ran the fraud check after committing.
The canonical trace teardown. One refund. Two paths. Same output, same JSON, same amount. One run checks the policy version, classifies the exception, and reviews fraud risk before committing. The other commits, then runs the checks. Output evaluation passes both. The behavior contract passes one and fails the other. The artifact pack ships the contract, the trace pair, and a 30-line CI gate you can adapt.
-
BW-2026-002
Type 01 · paper-to-practicecandidatepaper-to-practiceTrajectory-aware eval moved from research to vendor. Now what?
Google Vertex's Agent Evaluation ships trajectory grading as a first-class metric. What it catches; what its single-trajectory framing misses; how to layer a behavior contract on top of it without rebuilding the pipeline.
-
BW-2026-004
Type 04 · build artifactpublished2026-05-27Beast Mode — a behavioral auditor for your coding agent.
Raising agents is frustrating. Not because they fail. Because they pass. A 7-layer runtime that audits every Claude Code turn, scores it on 10 behavioral dimensions, and accumulates evidence in an append-only ledger. 20,906 entries. 76.3% Beast Index. Parallelism is the top failure. Costs ~$1/month. Installable in ten minutes.
-
BW-2026-005
Type 04 · build artifactsuperseded2026-05-29Skills 2.0 — making Anthropic's prose skills a Python contract. SUPERSEDED
Original case-file. The headline trust-score lift (0.25 → 0.53) was measured against a self-built metric. Three-arm blind judging (BW-2026-006) falsified the central claim. Kept live for the historical record.
-
BW-2026-006
Type 04 · build artifact (correction)published2026-05-29Skills 2.0 — falsified, partially fixed, then fixed.
Three pre-registered experiments. The first (EXP-008) falsified the central thesis under blind judging — the v0.5 runtime lost to a prompt-engineered checklist by 1 median point on framework fidelity. The second (EXP-009) fixed procedural compliance (3% → 100%) by making
@stepnarration mandatory but still trailed the baseline on framework fidelity. The third (EXP-010) added LLM case-specific elaboration on top of the runtime narration and won the original claim: mean framework fidelity 3.83 vs prompt-checklist 3.26 (+0.57, blind-judged, κ = 1.000). Skills 2.0 v1.0 = transformer + runtime + elaborator. Three modules, each empirically load-bearing. -
BW-2026-007
Type 04 · build artifact (boundary)published2026-05-29Skills 2.0 — the generalization boundary.
BW-2026-006 said three experiments converged on transformer + runtime + elaborator. That was true on one skill:
growth-accounting. EXP-011 tested generalization on two more skills with the same blind protocol. Both pre-registered falsifier clauses triggered. v0.7 works on quantitative-with-aggregates skills (growth-accounting) and does not work on qualitative-soft (psych-framework, −1.26 vs A2) or procedural-with-artifacts (abw-design, −1.23 vs A2). Plus a methodological lesson: the EXP-008 "v0.5 baseline" for stub-step skills was confounded — the subagent's prose was doing the framework application, not the runtime. The boundary, mapped honestly. v0.8 mechanism redesign next. -
BW-2026-003
Type 03 · failure patterndraftrequires public-safe evidenceThe "self-correcting agent" anti-pattern.
An agent that retries silently after a failed precondition produces a clean output and a wrong audit trail. Five deployments, same shape. Named, characterised, with the contract clause that catches it.
What you will not get
No "in this issue I'll cover" preamble. No reading-time labels. No "5 things I learned about AI agents this week." No paragraph asking you to share with a colleague. No AI-generated summary of someone else's paper. No sponsored placement. No daily email. The cadence is bi-weekly, and when the work does not support an issue, you get nothing that week. A missed issue is preferable to a thin one.
Where this fits
Behavior Watch is one of three timescales. The papers are the load-bearing arguments. The workbench is the reproducibility instrument. Behavior Watch is the continuous-practice surface: field notes, selected free experiments, and public-safe case files about delegation-grade agents. If the issue describes a private production problem your team has, the commercial route is Work With Us.
Still here?
Subscribe.
Same form as above. Free. Bi-weekly. One unsubscribe link in every issue.