Cheat sheetΒΆ

Eight subcommands. One of them spends tokens - everything else loads, re-scores, compares and renders for free, which is the whole point of keeping measuring and scoring apart.

uv run trysquare <cmd> python -m trysquare <cmd> uv tool install . → trysquare on PATH
1

Set up

before a single token
init [directory] free

Writes the skeleton of a new experiment - scenario.toml, prompt.md, hypothesis.md, and a trysquare.toml when none is found walking up. Refuses to overwrite anything. Default directory: here.

score.py is deliberately not written. Nothing runnable ships, so validate refuses the fresh skeleton by name until the validator is yours.
validate <scenario> free

Checks a scenario end to end with no output directory and no token: the file loads, every referenced path exists, [repos] has an entry, the thinking precondition holds.

--config
config file; defaults to the nearest trysquare.toml walking up
The same refusals run applies, shared with it: validate cannot pass what run would refuse. A missing pi on PATH is a note, not a failure.
2

Measure

the only command that spends
run <scenario> -o <dir> spends tokens

Measures a scenario. --output / -o is required and roots everything that writes: one directory per experiment.

--output, -o
Required. Directory every output is written under
--config
config file; defaults to the nearest trysquare.toml
--repetitions N
override, stamped into the directory name (→ ..._n3/)
--concurrency N
override, recorded in state.json and the synthesis header
--timeout N
override, recorded in state.json
--only CELL
restrict to these cells, repeatable; leaves the matrix incomplete, so no synthesis is written and --resume finishes it later
--group-by-cell / --no-group-by-cell
file runs under runs/<cell>/<id>/, or keep runs/<id>/ flat and opaque. Grouped by default, blind when a form validator says a human scores by hand
--resume
fill only what produced nothing - a cell added to the scenario since the ledger was written counts, having produced nothing by definition; a validator failure is re-scored instead, at no token cost. Refuses a cell whose declaration changed since it was measured
--extend
carry the runs of this same experiment measured at fewer repetitions - they are the very same runs - and measure only the difference. Implies --resume; the lower matrix is copied, not moved
--overwrite [CELL]
measure every run again, whatever is on disk. What a launch has always done, made typeable so the question below can be skipped. Given a cell name, repeatable, only those cells are measured again and every other run is kept - one variant edited and re-measured, not the matrix re-bought
--until-complete [N]
after a pass, relaunch the runs that produced nothing, at most N passes total (default 3). Never a re-measurement
--dry-run
show the whole plan and write nothing at all - not even the output directory
--no-progress
never draw the live bar; also TRYSQUARE_NO_PROGRESS=1
It refuses before spending
  • A missing brick. Every referenced path is checked first - a mistyped prompt path once became the literal string sent to the agent as its task.
  • A thinking mismatch when the scenario uses subagents.
  • An agent with no model, from neither its file nor an override.
Overrides are announced and stamped
! OVERRIDE: repetitions 10 -> 3
! OVERRIDE: concurrency  5 -> 10
  • What changes the measurement enters the directory name, so it cannot corrupt a published matrix.
  • What changes the load stays in the same directory but is recorded: it conditions the retries, and therefore every cost column.
The record scrolls, the bar is pinned
  ok  rule / off      412s  15234 in / 812 out
  !!  ticket / high   901s  timeout: no attempt
  ⏹ runs ━━━━╸━━━━  23/60  left 1h 55m

Estimate is throughput since launch, drawn on a terminal only. Relaunching the same experiment overwrites it - the archive of previous versions is git. On a terminal it asks first: the difference, everything, or abort. Piped, under --dry-run or with TRYSQUARE_NO_PROMPT=1 there is no question and nothing changes.

3

Read the results

no tokens, ever
render <scenario> -o <dir> free

Rebuilds the tables from stored measures. Measuring and scoring are separate: a scoring defect costs a render, not another hour of wall clock.

--repetitions N
which matrix to read
--reference CELL
a rendering choice, not a measurement: score another cell as the baseline → synthesis_ref-rule-off.md
--html
also export each archived session to a standalone page, in the run's own directory. Opt-in: ~0.3 s per session, and needs pi
synthesis.html is written whenever synthesis.md is, by any command. An archive must render offline in five years.
replay <dir> --scenario <file> free

Reconstitutes archived runs: clones the etalon at its tag, applies the archived diff.patch, writes a fresh context beside each tree. Takes an experiment directory or one run inside it.

--scenario
Required. Whose validators re-score the archive
--rescore
re-run the script validators, then rewrite measures.json, state.json and the synthesis
--config
config file
  • Never touched: usage, duration, attempts - facts about the run, not about the scoring.
  • A judge is never re-run; the archived verdict is reused, because replay exists on the promise that it costs nothing.
  • An empty run is left alone, and another scenario's directory is refused by name.
compare <left> <right> free

Compares two experiments side by side and prints every difference. A different model is a legitimate axis; it just has to be declared rather than hidden.

  • Hard refusal on different etalons - a different baseline means the two measures are not of the same thing.
  • Cost columns are set aside unless retries are near zero on both sides, with the counts shown.
  • Rates only, and no verdict: a certified gap needs both cells measured in one scenario.
watch <matrix dir> free

Follows a running matrix in a browser, on 127.0.0.1. Reads the directory a launch printed as its output, and writes nothing into it. The runs in flight come from live.json, which run refreshes every second.

--port <n>
default: a free one
--no-open
print the address, open nothing
Counts, never a verdict, until the matrix is complete: runs are interleaved by repetition, so every cell holds the same handful at any moment and an interval over four runs swings at each one that lands. The verdict stays synthesis.md's, and the page links to it once there is one.
form <scenario> -o <dir> free

Generates a manual scoring form - shuffled, with cell names withheld, the same blinding as the judge and for the same reason. The id-to-cell mapping lives in state.json.

--ingest <form.toml>
merge a filled form back in
TOML, not markdown with blanks, for a mechanical reason: an absent key is a metric not yet filled, so the file parses at any point. On ingest a manual value may fill a hole but never overwrite a measured one.
parity [bench measures.json] free --smoke: wall clock

Checks this harness against the bench it replaces, cheapest layer first. Neither tool is the reference - two computations are compared over the same archived material, and the material arbitrates. Exits non-zero when an exact layer disagrees.

--archive <dir>
the bench's archived run directories → adds layer 1
--scenario <file>
whose validators re-score the archive → adds layer 2. One clone of the etalon per run
--smoke <dir>
an experiment directory: layer 4's mechanical checks
--workdir <dir>
where sessions live, for the thinking check
--reference
default base
--criterion
default overflow
The four layers
layer 1strippingexactfrom archived sessions
layer 2scoringexactfrom tag + diff.patch
layer 3aggregationexactfrom published per-run rows
layer 4launchingnot comparableit samples
Layer 4 checks only what does not depend on the sample - above all that the thinking level each session recorded equals the level its cell declared, and that the model which answered is still the one the scenario names. It concludes nothing about any configuration, and says so.

On disk

A scenario, minimally

[scenario]
name = "rule-vs-ticket"
hypothesis = "hypothesis.md"  # before

[task]
repo = "my-repo"     # via [repos]
etalon = "etalon-v1"  #* a tag
prompt = "tickets/vague.md"

[agent]
provider = "ilaas"   #*
model = "gemma-4-31b" #*
thinking = "off"      #*

[protocol]
repetitions = 10        #*
concurrency = 5
timeout = 900

[axes]  # order fixes the table
context = ["nothing", "rule"]

[values.context.rule]   # the delta
context = "AGENTS.md"

[[validation]]
mode = "script"   # or "judge"
command = "score.py"
metrics = ["in_scope", "tests"]

[verdict]
criterion = "in_scope"
reference = { context = "nothing" }

#* mandatory and never inherited. Paths are relative to the scenario file. The baseline of an axis is its first value and needs no delta; every other value must declare one.

One directory per experiment

out/rule-vs-ticket_etalon-v1_ilaas_gemma-4-31b_n10/
  state.json      cells, runs, valid / empty / failed, attempts
  measures.json   one line per run
  synthesis.md    scores, costs, gaps, verdicts - when complete
  synthesis.html  the same, standalone
  runs/<cell>/<id>/ context, configuration, diff, validation
    session/*.jsonl   the agent's own trace, one per attempt
    session/*.html    on render --html

The name is the experiment's identity, which is why an override that changes the measurement lands in it. Relaunching the same experiment overwrites it: a timestamped directory per launch would be optional stopping through the back door.

The cell level goes away with --no-group-by-cell, and is absent by default when a form validator scores by hand: an id nobody can read is what keeps a human from knowing which configuration they are grading.

Exit codes

0Success. For parity, every checked layer holds.
1A refusal, a failed check, or a rendering error - the message says which.
2Usage error, or a missing required argument.

A refusal is exit 1 with a plain message, never a traceback. A rendering failure says explicitly that the measures are intact and only render needs rerunning.

A validator, in one line

score.py <context.json>
  → {"metrics": {…}, "reasons": {…}}
    on stdout, exit 0

A bool aggregates as a rate, a number as a median, anything else is diagnostic. A declared metric the validator omits makes the run invalid; extras are stored but carry no verdict.

The invariants

each one a defect that was paid for
□ A run counts only if it consumed tokens. A cut stream leaves turns that are real and empty: no rule broken, no file touched, no test failed - which a naive harness records as exemplary.
□ Nothing that changes a measurement may be inherited. Provider, model, thinking, etalon and repetitions live in the scenario. There are no environment variables anywhere in this tool.
□ A validator that could not judge never yields a verdict. A crash, a timeout, unreadable JSON or a missing declared metric makes the run invalid, not false.
□ Two states only: established or inconclusive. A gap is publishable if the 95% interval of its resampled difference excludes zero. The seed is fixed.
□ Repetitions are declared in advance, and a matrix is never rerun "to see". A resume may only relaunch runs that produced nothing; attempts are counted, so an abusive resume leaves a trace. More repetitions are added rather than substituted: --extend carries what a lower matrix measured, and says so in state.json and in the synthesis.
□ What the harness injects is excluded from scoring. Otherwise scope scoring counts our own configuration as the agent's work, and every configured cell drops to zero.
□ A judge is blind, and where it cannot be, the harness says so. When the treatment is the prompt, handing the judge the prompt reveals the cell.
□ Durations compare only within one matrix. Runs are interleaved across cells so they see the same provider load; across matrices they are not comparable.
A try square tells a joiner whether a joint is true. This one tells you whether a measured difference is. The reasoning behind every refusal: command line reference · the invariants · uv run trysquare <cmd> --help