Scenario schema¶
Every key of a scenario file, what it does, and whether it is required.
A scenario is validated entirely at load, before anything is spent. Every refusal below happens with no tokens consumed.
[scenario]¶
Key |
Required |
Meaning |
|---|---|---|
|
no |
Short name, used in the output directory. Defaults to the file stem. |
|
no |
One line, printed at launch and used as the synthesis heading. |
|
no |
Path to a file stating what is predicted and what would falsify it. Checked to exist. |
Tip
Write the hypothesis before measuring. A hypothesis written afterwards is a conclusion wearing a disguise, and the point of the file is to make a disappointing result publishable rather than quietly reframed.
[task]¶
Key |
Required |
Meaning |
|---|---|---|
|
yes |
A logical name, resolved by |
|
yes |
A git tag, or a commit written out in full (40 hex characters), cloned fresh
for each run. Never a branch or the working tree. A tag reads well and can be
moved by whoever owns the repository, after which two matrices report the same
etalon and measured different code; a commit cannot move, and puts what was
measured in the directory name. |
|
no |
The task given to the agent: inline text, or a path to a file. |
|
when |
The suite that decides the |
|
no |
Commands to run before the suite, in order. A failure here means nobody judged. |
|
no |
Path patterns for what running the task leaves behind and is not the agent’s work. |
repo being logical is what makes a scenario portable: it carries no author’s
directory layout. A value containing / or ~ is a mistake the schema does not
prevent but the config resolution will.
test_command is required as soon as any validator declares the tests metric, and
the refusal happens at load time. Required by the metric rather than by the section:
a scenario that measures prose has no suite to name, and demanding one would be
ceremony.
test_command = "node --test 'game/**/*.test.js'"
Declared and never detected, which is the same lesson as the mandatory keys above
wearing different clothes. The obvious detection is npm test, whose meaning is read
from package.json - a file inside the perimeter the measured agent may edit. Broken
code plus a test script of echo ok scores green, and nothing in the output says so.
A detected command hands the choice of how a run is measured to the agent being
measured.
A string, and it stays one. The file is what you would type, context.json carries the
same string, and split_command turns it into an argv wherever one is needed - the loader
that vets it and the base that runs it come to the same rule, so what loads is what runs.
shlex is the shell’s own word splitting, quotes included, so the rule is one every author
already knows - and a glob still works when the runner expands it itself, as node --test
does.
No shell ever runs it. A word that only means something to a shell - &&, |, ;,
a redirection - is therefore named and refused at load time, rather than reaching the
runner as an argument and failing where nobody can read it.
prepare¶
For the steps a suite needs before it can run:
[task]
prepare = ["npm ci"]
test_command = "npm test"
Separate from test_command because their failures mean different things, and the
difference is the one this whole tool is built around. A prepare that fails - no network,
a dependency that will not install - means nobody judged, so the metric is unjudged rather
than false. The suite failing is a measurement.
Conflated into one list, a broken network would score an agent red on a column that can
carry the scenario’s validity condition - “could not judge” filed as “worked badly”, one
level up.
Each entry is one command, under the same rules as test_command. A repository that
needs none is worth preferring: nothing to install is what makes a validation replayable
from a tag and a diff months later.
artefacts¶
What running the task leaves behind that nobody wrote:
[task]
test_command = "python3 -m unittest discover -s tests -t ."
artefacts = ["__pycache__"]
The defect
Without it, a matrix measured against a real provider scored in_scope = false in
every run of every cell. The only thing outside scope was __pycache__/*.pyc,
dropped by the agent running the declared suite to check its own fix. The criterion
saturated at zero, the gap the matrix existed to measure came out +0 pts, and six
paid runs concluded nothing.
Worse than noise, because it is not random: the runs that scored out of scope were the ones where the agent bothered to verify itself. A second matrix, whose agents did not run the suite, scored 3/3 on the same code.
Declared and never detected, the same lesson as test_command. A built-in list
would be a guess about somebody else’s language, and it would eventually hide a file an
agent really did write.
A pattern matches the whole path with shell globbing, so *.pyc catches
__pycache__/counter.pyc; or it matches any component of it, so __pycache__
catches that same file and tests/__pycache__/t.pyc without needing a *. A trailing
slash is accepted, since node_modules/ is how the directory is usually written.
It filters what a verdict rests on and never what is recorded. touched stays
complete in the context and in measures.json, because hiding a measurement is the
other dishonesty this tool refuses. A validator subtracts:
work = run.touched - run.artefacts
See Writing a validator.
[agent]¶
Key |
Required |
Meaning |
|---|---|---|
|
yes |
Provider name. Never inherited. |
|
yes |
Model id. Never inherited. |
|
yes |
Reasoning level. Never inherited, and never omitted. |
All three raise when absent:
[agent].provider is required in the scenario and is never inherited: a value that
is not declared is a value inherited from whoever runs the tool
Warning
thinking is mandatory even when you want the machine’s default. There is no way to
say “whatever the machine does”, because that is the defect that made a thinking cell
identical to its baseline in every published matrix.
If a scenario uses subagents, the harness additionally refuses to run when the
declared level differs from the machine’s defaultThinkingLevel. A subagent’s level
cannot be declared anywhere, so what cannot be controlled is verified instead. See
Troubleshooting.
[protocol]¶
Key |
Required |
Meaning |
|---|---|---|
|
yes |
Runs per cell. Declared in advance; part of the output directory name. |
|
no |
Parallel runs. Falls back to the config, then 5. |
|
no |
Seconds per run. Falls back to the config, then 900. |
|
no |
Retries while nothing has been produced. Falls back to 3. |
A plan carries its own load. concurrency and timeout condition the retry count and
therefore every cost column, so whatever their origin they are written into
state.json and printed in the synthesis header.
attempts is not optional stopping: a run that consumed no tokens produced no result,
so there is nothing to select between. A run that did produce something is never
retried, whatever its result.
Cells: [axes], [values], [variants]¶
A scenario needs at least one cell, from a grid, from named variants, or from both.
Grid¶
[axes]
context = ["nothing", "rule", "careful ticket"]
thinking = ["off", "high"]
[values.context.rule]
context = "AGENTS.md"
[values.context."careful ticket"]
prompt = "tickets/careful.md"
[values.thinking.high]
thinking = "high"
The cartesian product, named by joining the axis values with / -
rule / high. Declaration order of the axes fixes the order of the rendered
table: first axis in rows, second in columns.
[values.<axis>.<value>] holds only the delta from [agent] and [task].
Important
The first value of an axis is the baseline and declares no delta. Every other value must declare one:
axis 'context': value 'tickett' declares no delta. Only the first value of an
axis ('nothing') is the baseline. Deltas declared for this axis: ['rule',
'careful ticket']
Without that rule, a misspelling produces a cell identical to the baseline, published twice under two names, with nothing to reveal it.
Variants¶
[variants.nothing]
# the baseline: no delta
[variants."+subagents"]
harness = ["extension", "agents"]
Irregular cells, named explicitly. Grid and variants add rather than exclude, so a scenario may carry a regular grid plus a couple of named witnesses.
A cell name declared twice raises.
What a delta may contain¶
Key |
Effect |
|---|---|
|
Replaces the task prompt. Inline text or a path. |
|
Writes an |
|
Writes a |
|
Overrides the reasoning level for this cell. |
|
A list of brick names from |
[harness.<name>]¶
Bricks, declared once and cited by name so the pinning lives in one place and cannot diverge between cells.
[harness.extension] # a pinned repository
repo = "subagent" # logical name, resolved by [harness] in the config
tag = "formation-ai4dev-2026-v1"
load = "extension" # subdirectory passed to the agent
[harness.local] # a single file, relative to the scenario
load = "extensions/my-hook.ts"
[harness.agents] # files copied into the clone
paths = ["agents/explorer.md", "agents/tester.md"]
model = "ilaas/gemma-4-31b" # optional override, see below
[harness.skills]
paths = ["skills/profile"]
A brick with repo is cloned at its tag once per matrix, its dependencies are
installed, and every cell loads the same pinned state. An experiment that pins the
measured repository and lets the harness float is measuring the operator.
kind declares what a brick’s paths are - "skills" or "agents". Absent, the
brick named skills carries skills and every other brick carries agents. Declaring it
lets several skill bricks coexist, so variants compare skills one at a time instead of
loading one all-or-nothing brick:
[harness.skill-tdd]
kind = "skills"
paths = ["skills/tdd"]
[harness.skill-research]
kind = "skills"
paths = ["skills/research"]
[variants."+tdd"]
harness = ["skill-tdd"]
[variants."+research"]
harness = ["skill-research"]
[variants."+both"]
harness = ["skill-tdd", "skill-research"]
A kind outside the known values, or on a brick without paths, is refused at load
time: a misspelled kind would silently fall back to agents, and the cell would measure
subagents where its author declared skills.
The files kind¶
The third kind, and the only one whose material lands in the measured tree rather than in the agent library. It is how a cell hands the task a file the tag does not hold - a probe, a fixture, a specification the repository’s own test command will pick up:
[harness.probe]
kind = "files"
[harness.probe.files]
"game/probe.test.js" = "bricks/probe.test.js"
[variants."+probe given"]
harness = ["probe"]
A table of destination = source, and not a list of paths, because a list gives
the destination no name: the file would land under whatever basename it happens to
carry in the scenario’s directory, and the path a probe occupies decides whether
node --test 'game/**/*.test.js' finds it. That is a decision of the experiment.
Sources are relative to the scenario; destinations are relative to the clone and may
neither be absolute nor climb out of it.
Important
Given files are committed, not hidden.
Every other brick is written into .git/info/exclude, because harness plumbing is not
the agent’s work and scope scoring must not count it. A files brick is the exception:
its material is addressed to the task, and the agent may edit it or delete it.
Excluded, that would leave no trace anywhere - git ignores an untracked path however it
was left.
So the harness commits them on top of the etalon before the agent starts. The injection
still costs nothing in touched and nothing in diff.patch, since both are read
against HEAD; the moment the agent weakens a probe or deletes it, that shows up like
any other change. A replay puts the same files back, and commits them the same way,
before applying the patch.
A files brick never replaces what the tag holds. Overwriting a tracked file would
change the measured code while looking like nothing at all - the diff would be taken
against a HEAD that already contains the replacement. It is refused.
Note
[harness.agents].model is an optional override. Absent, each agent file declares its
own model, which keeps a cheap explorer alongside an expensive coder expressible.
Present, it overrides every file.
Either way the harness refuses to inject an agent that would end up with no model:
nine shipped agents once ran on the wrong provider and returned 402s because they
declared none. Two places may declare, so the trace settles which one applied -
configuration.json records the model used and where it came from.
Important
A cell that injects agents also loads the subagent gate, without declaring it.
Dropping agent definitions into the clone does not make them the only reachable ones.
The subagent tool takes its scope as a parameter the model chooses, and the default
reaches the agent library’s own built-in agents - none of which declares a model, so
each would inherit whatever the operator’s machine defaults to. A cell injecting
explorer would have measured someone else’s agent on someone else’s settings, and
nothing in the output would have said so.
trysquare/agent-gate.ts ships inside the package and is appended to the extensions
of any cell whose paths carry agents. It forces the scope and refuses
any agent name the scenario did not inject. It is not a [harness] entry, because
forgetting it was one line away from measuring the wrong thing.
[[validation]]¶
Repeatable. At least one is required - a scenario that measures nothing is refused.
Key |
Required |
Meaning |
|---|---|---|
|
yes |
|
|
yes |
The names this validator contracts to return. |
|
script |
Path to the executable. Always treated as a path. |
|
judge |
Path to the rubric. Always treated as a path. |
|
judge |
The judge’s own pinned configuration. |
|
judge |
What is assembled into the judge’s prompt: |
Two validators cannot declare the same metric. Refused at load, before any measurement:
metric declared by two validators: overflow. Rename one (for instance
overflow_judge)
See Writing a validator for the contract each mode implements.
[verdict]¶
Key |
Required |
Meaning |
|---|---|---|
|
yes |
The metric carrying the verdict. Must be declared by some validator. |
|
yes |
The cell every gap is measured against. |
|
no |
Metrics that must be true for a run to enter an aggregate. |
|
no |
Resampling draws. Default 10 000. |
|
no |
Resampling seed. Default 20260729. |
reference takes two forms because there are two ways to name a cell:
reference = { context = "nothing", thinking = "off" } # a grid
reference = "nothing" # a variant
A partial grid reference raises, as does a reference that is not a cell, and a criterion no validator declares.
Warning
validity must match the task. delivered requires a file to have changed - the
right condition for a task that asks for a diff, and exactly the wrong one for a task
that asks for prose. Getting it wrong eliminates the whole matrix, and the error names
which condition removed how many runs.