Cheat sheetΒΆ
Eight subcommands. One of them spends tokens - everything else loads, re-scores, compares and renders for free, which is the whole point of keeping measuring and scoring apart.
Set up
before a single token
Writes the skeleton of a new experiment - scenario.toml,
prompt.md, hypothesis.md, and a
trysquare.toml when none is found walking up.
Refuses to overwrite anything. Default directory: here.
score.py is deliberately not written. Nothing runnable ships,
so validate refuses the fresh skeleton by name until the validator is
yours.
Checks a scenario end to end with no output directory and no
token: the file loads, every referenced path exists,
[repos] has an entry, the thinking precondition holds.
- --config
- config file; defaults to the nearest
trysquare.tomlwalking up
run applies, shared with it: validate cannot
pass what run would refuse. A missing pi on PATH is a note, not
a failure.
Measure
the only command that spends
Measures a scenario. --output / -o is
required and roots everything that writes: one directory per
experiment.
- --output, -o
- Required. Directory every output is written under
- --config
- config file; defaults to the nearest
trysquare.toml - --repetitions N
- override, stamped into the directory name (→
..._n3/) - --concurrency N
- override, recorded in
state.jsonand the synthesis header - --timeout N
- override, recorded in
state.json - --only CELL
- restrict to these cells, repeatable; leaves the matrix incomplete, so
no synthesis is written and
--resumefinishes it later - --group-by-cell / --no-group-by-cell
- file runs under
runs/<cell>/<id>/, or keepruns/<id>/flat and opaque. Grouped by default, blind when aformvalidator says a human scores by hand - --resume
- fill only what produced nothing - a cell added to the scenario since the ledger was written counts, having produced nothing by definition; a validator failure is re-scored instead, at no token cost. Refuses a cell whose declaration changed since it was measured
- --extend
- carry the runs of this same experiment measured at fewer repetitions -
they are the very same runs - and measure only the difference. Implies
--resume; the lower matrix is copied, not moved - --overwrite [CELL]
- measure every run again, whatever is on disk. What a launch has always done, made typeable so the question below can be skipped. Given a cell name, repeatable, only those cells are measured again and every other run is kept - one variant edited and re-measured, not the matrix re-bought
- --until-complete [N]
- after a pass, relaunch the runs that produced nothing, at most N passes total (default 3). Never a re-measurement
- --dry-run
- show the whole plan and write nothing at all - not even the output directory
- --no-progress
- never draw the live bar; also
TRYSQUARE_NO_PROGRESS=1
- A missing brick. Every referenced path is checked first - a mistyped prompt path once became the literal string sent to the agent as its task.
- A thinking mismatch when the scenario uses subagents.
- An agent with no model, from neither its file nor an override.
! OVERRIDE: repetitions 10 -> 3 ! OVERRIDE: concurrency 5 -> 10
- What changes the measurement enters the directory name, so it cannot corrupt a published matrix.
- What changes the load stays in the same directory but is recorded: it conditions the retries, and therefore every cost column.
ok rule / off 412s 15234 in / 812 out
!! ticket / high 901s timeout: no attempt
⏹ runs ━━━━╸━━━━ 23/60 left 1h 55m
Estimate is throughput since launch, drawn on a terminal only. Relaunching the
same experiment overwrites it - the archive of previous
versions is git. On a terminal it asks first: the difference,
everything, or abort. Piped, under --dry-run or with
TRYSQUARE_NO_PROMPT=1 there is no question and nothing changes.
Read the results
no tokens, everRebuilds the tables from stored measures. Measuring and scoring are separate: a scoring defect costs a render, not another hour of wall clock.
- --repetitions N
- which matrix to read
- --reference CELL
- a rendering choice, not a measurement: score another cell as the
baseline →
synthesis_ref-rule-off.md - --html
- also export each archived session to a standalone page, in the run's own
directory. Opt-in: ~0.3 s per session, and needs
pi
synthesis.html is written whenever synthesis.md is, by
any command. An archive must render offline in five years.
Reconstitutes archived runs: clones the etalon at its tag, applies the archived
diff.patch, writes a fresh context beside each tree.
Takes an experiment directory or one run inside it.
- --scenario
- Required. Whose validators re-score the archive
- --rescore
- re-run the script validators, then rewrite
measures.json,state.jsonand the synthesis - --config
- config file
- Never touched:
usage,duration,attempts- facts about the run, not about the scoring. - A judge is never re-run; the archived verdict is reused, because replay exists on the promise that it costs nothing.
- An
emptyrun is left alone, and another scenario's directory is refused by name.
Compares two experiments side by side and prints every difference. A different model is a legitimate axis; it just has to be declared rather than hidden.
- Hard refusal on different etalons - a different baseline means the two measures are not of the same thing.
- Cost columns are set aside unless retries are near zero on both sides, with the counts shown.
- Rates only, and no verdict: a certified gap needs both cells measured in one scenario.
Follows a running matrix in a browser, on 127.0.0.1. Reads the
directory a launch printed as its output, and writes nothing into
it. The runs in flight come from live.json, which
run refreshes every second.
- --port <n>
- default: a free one
- --no-open
- print the address, open nothing
synthesis.md's, and the page links to it once there is one.
Generates a manual scoring form - shuffled, with cell names
withheld, the same blinding as the judge and for the same reason. The
id-to-cell mapping lives in state.json.
- --ingest <form.toml>
- merge a filled form back in
Checks this harness against the bench it replaces, cheapest layer first. Neither tool is the reference - two computations are compared over the same archived material, and the material arbitrates. Exits non-zero when an exact layer disagrees.
- --archive <dir>
- the bench's archived run directories → adds layer 1
- --scenario <file>
- whose validators re-score the archive → adds layer 2. One clone of the etalon per run
- --smoke <dir>
- an experiment directory: layer 4's mechanical checks
- --workdir <dir>
- where sessions live, for the thinking check
- --reference
- default
base - --criterion
- default
overflow
| layer 1 | stripping | exact | from archived sessions |
| layer 2 | scoring | exact | from tag + diff.patch |
| layer 3 | aggregation | exact | from published per-run rows |
| layer 4 | launching | not comparable | it samples |
On disk
A scenario, minimally
[scenario] name = "rule-vs-ticket" hypothesis = "hypothesis.md" # before [task] repo = "my-repo" # via [repos] etalon = "etalon-v1" #* a tag prompt = "tickets/vague.md" [agent] provider = "ilaas" #* model = "gemma-4-31b" #* thinking = "off" #* [protocol] repetitions = 10 #* concurrency = 5 timeout = 900 [axes] # order fixes the table context = ["nothing", "rule"] [values.context.rule] # the delta context = "AGENTS.md" [[validation]] mode = "script" # or "judge" command = "score.py" metrics = ["in_scope", "tests"] [verdict] criterion = "in_scope" reference = { context = "nothing" }
#* mandatory and never inherited. Paths are relative
to the scenario file. The baseline of an axis is its first value
and needs no delta; every other value must declare one.
One directory per experiment
out/rule-vs-ticket_etalon-v1_ilaas_gemma-4-31b_n10/ state.json cells, runs, valid / empty / failed, attempts measures.json one line per run synthesis.md scores, costs, gaps, verdicts - when complete synthesis.html the same, standalone runs/<cell>/<id>/ context, configuration, diff, validation session/*.jsonl the agent's own trace, one per attempt session/*.html on render --html
The name is the experiment's identity, which is why an override that changes the measurement lands in it. Relaunching the same experiment overwrites it: a timestamped directory per launch would be optional stopping through the back door.
The cell level goes away with --no-group-by-cell, and is absent by
default when a form validator scores by hand: an id nobody can read
is what keeps a human from knowing which configuration they are grading.
Exit codes
| 0 | Success. For parity, every checked layer holds. |
| 1 | A refusal, a failed check, or a rendering error - the message says which. |
| 2 | Usage error, or a missing required argument. |
A refusal is exit 1 with a plain message, never a traceback. A
rendering failure says explicitly that the measures are intact and only
render needs rerunning.
A validator, in one line
score.py <context.json>
→ {"metrics": {…}, "reasons": {…}}
on stdout, exit 0
A bool aggregates as a rate, a number as a median, anything else is diagnostic. A declared metric the validator omits makes the run invalid; extras are stored but carry no verdict.
The invariants
each one a defect that was paid for--extend carries what a lower matrix measured, and says so
in state.json and in the synthesis.