trysquare¶
A try square tells a joiner whether a joint is true. This one tells you whether a measured difference is.
A scenario harness for measuring coding agents reproducibly.
One scenario is one self-contained experiment in one TOML file: the task, the configurations to compare, the protocol, and the validation. The harness runs it, scores it, and refuses to publish a difference that does not survive resampling.
Why this exists
Measuring an agent is easy to get wrong in ways that look right.
The tool this one replaces produced six published conclusions that collapsed on rerun, and twenty catalogued defects. Most were variations on a single mistake: a run that did not do the work reading as a run that worked well. A provider cuts the stream, the agent retries and returns turns that are real but empty. That run breaks no rule, touches no file, and fails no test - so a naive harness records it as exemplary.
Every rule in this tool is a defect that was paid for. The documentation says which one, because a rule whose reason is lost gets removed by the next person who finds it inconvenient.
Start here¶
Install, write a skeleton, plan a run for free, measure, read the output.
Scenario, cell, brick, validator, verdict. The vocabulary the rest assumes.
Grids, variants, bricks, protocol. Every key explained.
The eight rules that make a number publishable, and the defect behind each.
Every command, every flag, and what it costs, on one page.
Reference¶
User guide
- Getting started
- Core concepts
- Writing a scenario
- Writing a validator
- The invariants
- 1. A run counts only if it consumed tokens
- 2. Nothing that changes a measurement may be inherited
- 3. A validator that could not judge never yields a verdict
- 4. Two verdict states, and only two
- 5. Repetitions are declared in advance
- 6. What the harness injects is excluded from scoring
- 7. A judge is blind, and where it cannot be, the harness says so
- 8. Durations compare only within one matrix
- 9. A run that was cut short is not recorded at all
- And one that is not a rule but a habit
- Parity with the previous bench
- Troubleshooting
- “these files the scenario references do not exist”
- “init never overwrites”
- “refused: the scenario declares thinking = … “
- “these agents declare no model”
- “metric declared by two validators”
- “value ‘x’ declares no delta”
- “reference cell … has no run left to compare against”
- “
piis not on PATH” - Runs marked
empty - A validator failed
- “this scenario names …, and you asked to re-score …”
- A validator refuses “the context carries no …”
- The synthesis warns about cost columns
- No synthesis was written
- “
--onlynames no cell of this scenario” - “these cells changed since their runs were measured”
- A harness clone failed
- A repository URL could not be cloned
- A resume against a URL wants the network again
- “[repos] has no entry”
- “no trysquare.toml was found”
- A repository path does not exist
- “refused: different etalons”
- No progress bar appears
- The cursor is gone after a hard kill
- Stopping a matrix
In one page¶
pip install -e . # or: uv sync
cp trysquare.toml my-trysquare.toml # point [repos] at your repository
uv run trysquare run my-scenario.toml -o out --dry-run # plan, spend nothing
uv run trysquare run my-scenario.toml -o out # measure
[scenario]
name = "rule-vs-ticket"
[task]
repo = "my-repo" # a logical name, resolved by the config file
etalon = "etalon-v1" # a tag, cloned; never the working tree
prompt = "tickets/vague.md" # relative to the scenario; inline text works too
[agent]
provider = "ilaas" # mandatory, never inherited
model = "gemma-4-31b" # mandatory
thinking = "off" # mandatory
[protocol]
repetitions = 10 # declared in advance
concurrency = 5
timeout = 900
[axes] # a grid; declaration order fixes the table's order
context = ["nothing", "rule"]
thinking = ["off", "high"]
[values.context.rule]
context = "AGENTS.md"
[values.thinking.high]
thinking = "high"
[[validation]]
mode = "script"
command = "score.py"
metrics = ["in_scope", "delivered", "tests"]
[verdict]
criterion = "in_scope"
reference = { context = "nothing", thinking = "off" }
validity = ["delivered", "tests"]
Requirements¶
Python >= 3.11: TOML parsing is tomllib from the standard library, which is why
the floor is there. uv sync - or pip install -e . - installs the rest.
Note
Measuring anything also needs the agent binary (pi) on PATH and a provider you
have access to. Everything else - loading, scoring, aggregation, verdicts, the parity
checks - runs offline, which is why the methodological rules have tests at all.
Status¶
Working and tested, with gaps that are listed rather than left to be discovered. See Outputs for what is produced today and Not implemented yet for what is not.