Training Agents
Raising Agents

Behavior Watch newsletter · delegation-grade agents · public-safe lab

Adoption is solved.
Value is not.

Most organizations can get agents to produce plausible work. They still hesitate to delegate significant work, because plausible output is not permission to act. Raising Agents studies the missing layer: described work, authority boundaries, evidence, thresholds, receipts, and review loops that let agents earn delegation.

The frame: describe -> delegate -> trust -> autonomy -> value. You cannot delegate what you cannot describe.

The center is Behavior Watch: the newsletter where field notes, selected experiments, trace teardowns, and general lessons can travel safely without turning private Zartis work, client work, or non-shareable project IP into content.

Raising Agents is a personal research initiative by Adrian Sanchez de la Sierra, developed alongside his AI and innovation work at Zartis. It is independently written, while many of the experiments and infrastructure are made possible by Zartis' investment in applied AI research. It is not an official Zartis publication unless explicitly marked. Research and Lab pages publish only repos, experiments, artifacts, and examples that can be shared: public, synthetic, redacted, open-source, or explicitly approved.


Proof instrument

Would you let an agent do this?

Give Raising Agents a task. It will separate the answer from the permission to act: the hidden action, the evidence required, the authority owner, the contract boundary, and whether the agent can proceed, clarify, escalate, or stop.

Public, synthetic, redacted, or generalized examples only.


The pain

AI is everywhere. Value is not.

The adoption threshold has been crossed. The value threshold has not. The useful question is no longer only whether a model can answer, reason, or call tools. It is whether an organization can safely authorize delegated work when the answer looks plausible.

/ summit frame

88%

of organizations have adopted AI in the talk framing. Adoption is not the hard part anymore.

/ value capture

6%

capture meaningful value. The gap persists because plausible work is not yet delegated, verified, authorized, and learned from.

/ root cause

0%

cite model quality in the summit root-cause split. The stronger explanation is the system around delegated work, not the model alone.

The model produces plausible work. The organization cannot yet authorize it.

That is why significant work stays with humans. Not because agents cannot produce anything useful, but because most workflows lack explicit owners, proceed conditions, evidence requirements, escalation paths, rollback paths, and receipts. The delegation threshold is invisible.


Behavior Watch

The field journal for delegation-grade agents.

Essays, trace teardowns, failure patterns, selected experiment results, tool notes, and public-safe reconstructions. Each issue asks the same question: what would make this agent safe to delegate significant work to?


Free Experiments

Give the lab a serious agent behavior problem. If it is sharp and public-safe, we may run the experiment for free.

The intake is an agent interview, not a static form. It pushes for the failure pattern, why normal evals miss it, what a public-safe reconstruction could test, and what evidence would falsify the idea. Selected ideas can become Lab experiments or Behavior Watch case files.


Live lab desk

The queue is part of the product.

This is the public edge of the workbench: current experiment ideas, stuck points, shipping notes, and reader signals. Add a problem and the Free Experiments agent will turn it into a sharper candidate.

Public, synthetic, redacted, or generalized examples only. The full interview opens on Free Experiments.


Shareable evidence

Why output review is not enough.

An AI agent can return the right refund, the right claim, or the right edit through the wrong process. The eval sees the output and passes. The trace records whether the work was actually safe to delegate.

/ EXP-002 · θ_OPBR

91.4%

of behavioral regressions produced passing outputs across four domains. Output-only evaluation misses the class of failure Raising Agents studies.

/ EXP-004 · held-out

F1 = 0.982

contract detection on held-out OPBR. The best rich-output provenance baseline reached F1 = 0.400.

/ EXP-001 · controlled

180/180

induced regression runs passed output evaluation while behavior contracts failed. The behavior layer separated the traces.

Open the research hub →

Delegation stack

Delegation-grade agents need a system around the model.

The research stack is built from small public-safe components: describe the work, govern authority, run the procedure, protect the action boundary, record evidence, and verify claims. OPBR is one empirical proof. The larger problem is delegation control.

Charter
ratified principles, authority tables, rubrics, and policy rails.
Writ
typed, ratified procedures for recurring delegated work.
AAA
single-use action leases before consequential external effects.
Attest
flight recorder and claim-evidence gate for agent work.
Keel
local-first evidence and decision control plane.

Routes

/01

Watch

Behavior Watch is the center: field notes, case files, trace notes, experiment results, and general lessons that can travel without exposing private project IP.

Subscribe →

/02

Free Experiments

An agent-led intake for serious agent behavior problems. Selected public-safe ideas may become free lab experiments or Behavior Watch cases.

Submit →

/03

Research

The evidence hub for papers, lab experiments, schemas, methods, workbench tools, and repos that are public, open-source, redacted, synthetic, or explicitly approved.

Read →

/04

Work With Us

Start with Adrian for advisory. Route larger, private, implementation-heavy, or team-specific work through Zartis.

Choose route →

Boundary

The newsletter can explain the pattern. Free Experiments and the Lab only share what can be shared.

Open
Newsletter essays, public repos, papers, lab protocols, synthetic examples, redacted examples, and approved artifacts.
Zartis-routed
Workshops, assessments, implementation support, production adoption, and team-specific conversations.
Blocked
Client traces, confidential Zartis IP, private transcripts, credentials, non-shareable repos, and engagement-specific detail.