Command line reference¶
uv run python -m trysquare <command> [options]
trysquare <command> [options] # if installed
--output roots everything that writes. Eight commands.
init¶
Writes the skeleton of a new experiment - scenario.toml, prompt.md,
hypothesis.md, and a trysquare.toml when none is found walking up - into a
directory, refusing to overwrite anything. The validator is deliberately not
written: nothing runnable ships, and validate refuses the fresh skeleton by
name until score.py is yours.
trysquare init [directory] # default: here
run¶
Measures a scenario.
trysquare run <scenario> --output <dir> [options]
Option |
Effect |
|---|---|
|
Required. Directory every output is written under. |
|
Config file. Defaults to the nearest |
|
Override. Stamped into the directory name. |
|
Override. Recorded in |
|
Override. Recorded in |
|
Restrict to these cells. Repeatable. Leaves the matrix incomplete. |
|
File each run under a directory named for its cell, or keep |
|
Fill only what produced nothing. |
|
Carry the runs of this same experiment measured at fewer repetitions, and measure
only the difference. Implies |
|
Measure every run again, whatever is on disk. The default, made typeable. Given a cell name, repeatable, only those cells are measured again and every other run is kept. |
|
After a pass, relaunch the runs that produced nothing, at most N passes in
total (default 3). Never a re-measurement: a run that produced something is
out of reach of every pass, exactly as it is for |
|
Show the plan and write nothing at all. |
|
Never draw the live bar. Also |
The record scrolls, the bar is pinned¶
ok rule / off 412s 15234 in / 812 out 4 turns 0 retries
!! careful ticket / high 901s ... timeout: no productive attempt
⠹ runs ━━━━━━━━━━╸━━━━━━━━━━ 23/60 elapsed 1h 12m left 1h 55m
Every per-run line still prints, unchanged, above the bar: those lines are the record, and the bar is not. What the bar adds is the part a terminal never said - how many runs are left, and when the matrix is expected to land.
The estimate is the throughput since launch, elapsed / completed x remaining, not
a recent rate. Runs finish in batches the width of --concurrency, so a sliding-window
estimate would swing by minutes every few seconds while the true arrival time barely
moves. Nothing is claimed until the first run lands.
The bar is drawn only on a terminal. Piped, redirected, or captured by a test, the
output is exactly the bytes it was before the bar existed - no escape sequences in a log
file. --no-progress, TRYSQUARE_NO_PROGRESS=1 and TERM=dumb each turn it off on a
terminal too. NO_COLOR does not: it asks for no colour, not for no motion.
Overrides are announced and stamped¶
! OVERRIDE: repetitions 10 -> 3
! OVERRIDE: concurrency 5 -> 10
Every override is printed at launch. The previous tool had a protocol declared in a document and defaults in the code that contradicted it - and the code wins at the moment somebody types the command, so a published matrix was measured at the wrong load. A plan that cannot be executed is not a plan.
Then, according to effect:
What changes the measurement enters the directory name. --repetitions 3 writes
to ..._n3/, so it cannot corrupt a published matrix at ten. This is not a check
bolted on; the name already carries the experiment’s identity.
What changes the load stays in the same directory but is recorded, because it conditions the retry count and therefore every cost column.
--only leaves the matrix incomplete, so no synthesis is written and --resume
completes it later. The two compose: a cell measured alone for debugging leaves a
resumable matrix rather than a dead end.
Measuring one variant again¶
An author who edits one variant of a matrix that is already measured wants that variant
re-measured, and not the other fifty runs re-bought. --overwrite takes the cells to
measure again, repeatably:
$ trysquare run scenario.toml -o out --overwrite "careful ticket / high"
! REPLAY: 10 runs of careful ticket / high are measured again; the 50 runs of the
other cells of rule-vs-ticket_..._n10 are kept
10 runs to perform
Kept means kept whole: the ledger entry, the row in measures.json and the run tree of
every other cell are left exactly as they were found. A replay restricts the launch
the way --only does - runs of other cells that produced nothing are left alone too, and
a note says so - but it does not declare the matrix partial. Once the replayed runs land
and nothing else is missing, the matrix is whole and the synthesis is written, which is
the point of measuring one variant again.
The two lists cannot both be given: --overwrite CELL already says which cells a launch
restricts itself to, so --only beside it would say it twice and is refused.
A cell that changed under its own name is still refused - for the cells a replay keeps, since publishing their old results beside the new ones is the defect the fingerprint exists to catch. The refusal names the replay that resolves it. The cells being measured again are free to declare something new: that is what they are for.
When results already exist, it asks¶
$ trysquare run scenario.toml -o out
rule-vs-ticket_..._n10 already holds 54 runs that produced a result and 6 that
produced nothing.
[d] difference measure only those 6 (--resume)
[e] everything measure them all again, resetting this ledger (--overwrite)
[a] abort nothing is spent
>
Asking for more repetitions than a matrix on disk asks the other flavour of the same
question: carry the lower matrix’s runs and measure the difference (--extend), measure
everything from scratch, or stop.
The question is asked once, before the plan is made. --until-complete resolves a
plan per pass, so a question further in would be asked once per pass.
A flag is never second-guessed. --resume, --extend and --overwrite are the three
answers; give one and nothing is asked. They are mutually exclusive, because giving two
says two things at once.
Off a terminal there is no question. Piped, redirected, under --dry-run, with
TRYSQUARE_NO_PROMPT=1 or TERM=dumb, a launch does exactly what it did before the
question existed: the announced overwrite. That is what keeps a scripted matrix
reproducible - and a prompt written into a log file with nobody to answer it would hang a
pipeline on an invisible question. Both streams have to be terminals for that reason.
There is no default answer, and Enter is not one: the cheapest keystroke must not be the one that spends a matrix of tokens. Aborting exits 1 - a launch that measured nothing must not report success to whatever ran it.
More repetitions, without paying twice¶
A run id is blake2s(scenario/cell/repetition) and ignores the repetition count, so
the sixty runs of a matrix at ten repetitions carry the very ids a matrix at twenty asks
for at its first ten. They are those runs, not analogues of them - which is what makes
carrying them a copy rather than a claim.
$ trysquare run scenario.toml -o out --repetitions 20
! AVAILABLE: rule-vs-ticket_..._n10 holds 60 measured runs this matrix asks for at
its first 10 repetitions. --extend carries them over and measures only the
difference; without it they are measured again
120 runs to perform
$ trysquare run scenario.toml -o out --repetitions 20 --extend
! CARRIED: 60 runs measured in rule-vs-ticket_..._n10 at 10 repetitions (concurrency
5, timeout 900s) are this matrix's own runs, carried and never re-measured. The
synthesis says so too
60 runs to perform
Nothing is carried silently. The record goes into state.json as a list, so a
matrix extended twice can be read back to the launch that first measured each run; the
note above is derived from that ledger rather than from the flag, so it prints on every
later launch too; and the synthesis carries both a header line and a :warning:
paragraph, for a reader who never saw the terminal.
Carried runs are filed under this matrix’s layout, whichever one they came from: the two layouts do not merge, so a blind matrix’s runs cannot arrive in a grouped one still flat. See the two layouts.
The lower matrix is copied, never moved. It stays publishable on its own, and
compare puts the two readings side by side. A link would be worse than a copy:
replay --rescore rewrites validation/<mode>.json in place, so re-scoring the
extended matrix would silently rewrite the one it came from.
Refused rather than quietly useless. --extend with no lower matrix names what is
under --output instead of measuring from scratch, for the reason --only with a typo
is refused. And a source that disagrees on thinking or on repo is refused by name:
neither is in a directory name, so two matrices can share a name without sharing what
they measured.
Only into a matrix that holds nothing yet. A carry is how a matrix begins as the
extension of a lower one. Once it has a ledger of its own, --extend is simply the
--resume it implies, and nothing is re-imported.
What it does not do is make the two halves comparable: runs are interleaved so that the durations of one matrix compare between its cells under one provider load, and an extended matrix holds two launches. That is written into the synthesis rather than argued away.
--dry-run writes nothing¶
Not even the output directory. That was once false - resolve() performed side
effects belonging to execute(), so a dry run against an existing experiment reset
its ledger.
Before spending anything¶
run refuses on three preconditions:
Missing bricks. Every referenced path is checked first.
Thinking mismatch when the scenario uses subagents.
An agent with no model, from neither its file nor an override.
The plan header states, before anything is spent: a duration bound (runs,
concurrency and timeout are declared, so runs / concurrency x timeout is
arithmetic, not a guess), a spend estimate from this experiment’s archive, an
OVERWRITE note when the output directory already holds a ledger, and - on
--dry-run - whether a real run would refuse for want of pi.
The spend estimate is the median of what the archived valid runs actually cost, times the runs to perform. Never a price list: the harness is provider-agnostic, a maintained price table goes stale silently, and a wrong estimate is worse than none.
Agents that report no price are common, and an archive nobody priced is not an empty archive - so it is used in the unit the provider does report, median tokens per run and what the plan comes to. Only a scenario that has never run says there is nothing to estimate from.
validate¶
Checks a scenario end to end without an output directory, and without spending a
token: the file loads, every referenced path exists, the config has an entry for
the repository, the thinking precondition holds. The same refusals run applies,
shared with it - validate cannot pass what run would refuse. A note (not a
failure) says when pi is missing from PATH, since a run would then refuse.
trysquare validate <scenario> [--config <file>]
render¶
Rebuilds tables from stored measures. Costs nothing.
trysquare render <scenario> -o <dir> [--repetitions N] [--reference CELL] [--html]
[--no-progress]
Measuring and scoring are separate. When you find a scoring defect - and there were three in one day on the previous tool - it would be absurd to pay for another hour of wall clock to fix it. Per-run values are persisted for exactly this.
--reference writes a suffixed file from the same measures:
trysquare render my-scenario.toml -o out --reference "rule / off"
# -> synthesis_ref-rule-off.md
A reference is a rendering choice, not a measurement. Changing it is how an
interaction is read without paying twice, and it is why the previous tool’s
hand-renamed _ref-thinking file was a symptom rather than a solution.
--html rebuilds the session pages¶
trysquare render my-scenario.toml -o out --html
# runs/658df337/session/2026-07-30T07-27-59-046Z_019fb1ec.html
# ...
# 12 session pages written
One standalone page per archived session, written in the run’s own directory beside
the jsonl it came from. The rendering is pi --export, so it costs no tokens and needs no
network - see session/*.html, on render --html.
Opt-in, because it is the only part of render that spends wall clock: roughly 0.3 s per
session, and a matrix at ten repetitions holds sixty of them.
It runs before the table and independently of it. No synthesis is written for an incomplete matrix, and an incomplete matrix is exactly when a trace is wanted, so a missing table must not take the pages down with it.
Two things are said rather than left to be discovered:
a run with no archived session is counted and named - an output tree measured before sessions were archived would otherwise answer with a silence that reads as success;
a session that will not render is reported on stderr and skipped. One broken trace does not cost the rest, the same rule that applies to a run inside a matrix.
Without pi on PATH the command refuses with a message and exit 1, rather than writing
nothing and claiming success.
Whenever synthesis.md is written - by run, render or replay --rescore -
synthesis.html is written beside it: the same synthesis as one self-contained
page, no script and no external asset, linking each run’s session pages when
render --html has produced them. An archive must render offline in five years.
replay¶
Reconstitutes archived runs so they can be re-scored. Costs no tokens.
trysquare replay <experiment dir or run dir> --scenario <scenario> [--config <file>]
trysquare replay <experiment dir> --scenario <scenario> --rescore [--no-progress]
Clones the etalon at its tag, applies the archived diff.patch, and writes a fresh
context beside each reconstituted tree, under <workdir>/replay/<run id>/. This is what
makes “fix a signature and re-score runs already paid for” true rather than aspirational:
the archive keeps the tag and the patch, not 150 copies of a working tree.
The context is written fresh rather than reused. The archived one holds absolute paths
into the work directory of the original run, which lives under $TMPDIR by default and
which the system may long since have purged - so it names a tree that no longer exists,
while the tree just rebuilt would be named by nothing. touched is recomputed from that
tree, files from the tag, and the session is the archived one.
--rescore¶
Re-runs the scenario’s script validators over the reconstituted trees, then rewrites
measures.json, state.json and the synthesis. Still no tokens.
This is what closes the loop. Without it, a corrected signature could be executed against
sixty archived runs and the sixty answers had nowhere to go: only run and form ever
wrote measures.json, so a reader had to assemble a second scoring path of their own -
which is exactly what a single harness exists to prevent.
What it may and may not touch:
usage,duration,attemptsnever. They are facts about the run, not about the scoring. In particular
attemptsis what leaves an abusive resume visible instate.json, and a re-scoring must not spend that.state.jsonrewritten with the new per-run state, and it has to be. That file decides whether a synthesis is publishable at all, so a validator repaired here would otherwise stay counted among the failures and the matrix would keep refusing to publish.
validation/<mode>.jsonreplaced by what the validator just returned. Leaving the old payload would put a
measures.jsonand avalidation/script.jsonside by side that disagree, with nothing to say which one is the score.
A judge is not re-run: its verdict costs tokens, and replay exists on the promise
that it costs none. The archived payload is reused instead, which is the right answer and
not a compromise - that verdict is a measurement somebody paid for, and correcting a script
metric must not silently discard it. A run whose archived judge verdict is missing is left
alone rather than scored short.
The same principle applied to a script metric that a replay cannot answer: one that had a value and no longer has one is named, one line per metric.
6 of 6 runs re-scored, no tokens spent
! documented no longer has a value on 6 of 6 runs: the context carries no 'response',
so the agent's final prose cannot be read
a measurement that was paid for is gone from the tables, and the
previous measures.json is in git - this archive's own history
Named rather than refused: re-scoring a metric of process out of an archive is a legitimate thing to do knowingly, and refusing would make the flag useless for the case it exists for. What is not acceptable is doing it in silence - a metric that becomes unjudged simply drops out of the score table, so the only sign used to be a column that stopped appearing.
Two refusals, both before anything is written:
a run whose state is
emptyis left alone. No scoring turns “produced nothing” into a measurement, and overwriting its state would hide an incomplete matrix behind a full-looking one.a directory that is not this scenario’s is refused by name. A directory name is the experiment’s identity, so comparing the two names catches a re-scoring that would rewrite measures belonging to another matrix.
Note
Three things a replay cannot give back: prompt, response and trace. The first two
lived in the work directory, and the raw stream is deliberately never archived. A
validator reading one of them refuses by name - “the context carries no response” -
rather than scoring a run it could not see. That is also why there is no context version
number: a named absence says more than a version could.
A metric of process does replay, which is worth knowing because it looks as though it should not: the agent’s tool calls are in the archived session, not only in the discarded stream.
Note
--rescore re-runs the validators one run at a time, in order. It is deliberately not
concurrent: a re-scoring is cheap, and the contexts live in per-run directories precisely
so that it could be, but nothing here needs it yet.
compare¶
Compares two experiments, refusing what is not comparable.
trysquare compare <left dir> <right dir>
Hard refusal on different etalons - a different baseline means the two measures are not of the same thing.
Cost columns set aside unless retries are near zero on both sides, with the counts shown.
Everything else is allowed and every difference is printed. A different model is a legitimate comparison axis; it just has to be declared rather than hidden.
When both sides hold measures, a side-by-side score table follows: one row per
cell, x/n -> y/n per boolean metric, a cell absent on one side named with -
rather than dropped. Rates only, and no verdict: costs are never compared across
matrices, and a certified gap needs both cells measured in one scenario.
parity¶
Checks this harness against the tool it replaces. See Parity with the previous bench.
trysquare parity <bench measures.json> [--archive <traces dir>]
trysquare parity <bench measures.json> --archive <traces dir> --scenario <scenario>
trysquare parity --smoke <experiment dir> [--workdir <dir>]
Layers run cheapest first: layer 3 always, layer 1 with --archive, layer 2 with
--scenario as well. With --smoke, layer 4’s mechanical checks over a matrix you
measured. The command exits non-zero when an exact layer disagrees.
--scenario is what layer 2 re-scores with: each archived run is reconstituted from the
tag and its diff.patch, then that scenario’s script validators run against the
tree. A judge is not run - its verdict costs tokens, and layer 2 costs none - so a metric
only a judge can produce is named as out of scope rather than reported as a
disagreement. Expect one clone of the etalon per run.
form¶
Generates or ingests a blind manual scoring form.
trysquare form <scenario> -o <dir> # generate
trysquare form <scenario> -o <dir> --ingest <form.toml>
The form is shuffled, with cell names withheld, the same blinding as the judge and
for the same reason: someone who knows they are scoring the best-equipped cell scores
it better. The id-to-cell mapping lives in state.json, deliberately not in the form.
[run.a7f3]
diff = "runs/a7f3/diff.patch"
# readable_in_class =
TOML rather than markdown with blanks, for a mechanical reason: an absent key is a
metric not yet filled, so the file parses at any point without a bespoke parser. Every
prose line is a comment for the same reason - --ingest reads this file back.
A tree grouped by cell puts the cell in every one of those paths, so
the form says so at the top of the file and form says so on the terminal:
! runs are grouped by cell, so this scenario's form names the configuration in every
path: whoever fills it can see which cell they are scoring
Not a refusal - grouping is asked for on purpose, and reading a matrix before anybody scores it is when it helps. But a form that names the cells is not a blind form, and it may not read as one.
On ingest, a manual metric may fill a hole but never overwrite a measured value:
2 manual metrics merged, 38 still pending, 1 refused
refused values would have overwritten a measured metric, which a form may never do
Exit codes¶
Code |
Meaning |
|---|---|
|
Success. For |
|
A refusal, a failed check, or a rendering error. The message says which. |
|
Usage error, or a missing required argument. |
A refusal is exit 1 with a plain message, never a traceback. A rendering failure says
explicitly that the measures are intact and only render needs rerunning.