Python API

The package is importable, and the split below is the architecture: the pure half never touches the world, which is why the methodological invariants have tests that need no network, no clone, and no API key.

The pure core

trysquare.scenario

Loading a scenario, and expanding it into the cells it describes.

A scenario is one self-contained experiment in one TOML file: the task, the configurations to compare, the protocol, and the validation. This module turns that file into data and refuses it when it is not an experiment.

Nothing here touches the disk beyond reading the file it is handed, and nothing here runs an agent. That is deliberate: every rule below is a rule about what counts as a well-formed experiment, and those rules are worth testing without a network, a clone, or an API key.

trysquare.scenario.split_command(command)[source]

One declared command, as an argv.

The single splitting rule, and the only one: the loader vetting a scenario and the base running the command both come here. Two implementations would be the drift this module spends its whole length refusing - and a command split two slightly different ways would be measured two slightly different ways.

shlex is the shell’s own word splitting, quotes included, which is why the scenario can carry a string an author would recognise. Nothing here runs a shell; SHELL_ONLY above is what makes that safe.

Parameters:

command (str)

Return type:

tuple[str, …]

exception trysquare.scenario.ScenarioError[source]

Bases: Exception

A scenario that is not a well-formed experiment.

Always raised at load time, never at measure time: the point is to fail before spending tokens, not after.

class trysquare.scenario.Cell(name, delta=<factory>, description='')[source]

Bases: object

One configuration to measure, and how it differs from the baseline.

description is prose the scenario may declare about why this cell exists, and it is deliberately not in delta. The delta is what the cell changes, which is what outputs.cell_fingerprint hashes and what a reader can check against the run; a sentence about it is neither. Kept apart, editing the prose cannot make a resume refuse the runs it describes, and is_baseline still reads an empty delta as the baseline rather than as a cell that explains itself.

Parameters:
name: str
delta: dict
description: str = ''
property is_baseline: bool
class trysquare.scenario.Validator(mode, metrics, config=<factory>)[source]

Bases: object

One validator, and the metrics it contracts to return.

metrics is a contract rather than documentation: a run whose validator omits one of these names is invalid, and a metric nobody declared can never carry a verdict.

Parameters:
mode: str
metrics: tuple[str, ...]
config: dict
class trysquare.scenario.Scenario(name: 'str', title: 'str', task: 'dict', agent: 'dict', protocol: 'dict', cells: 'tuple[Cell, ...]', validators: 'tuple[Validator, ...]', verdict: 'dict', bricks: 'dict' = <factory>, hypothesis: 'str | None' = None, axes: 'dict' = <factory>, path: 'Path | None' = None)[source]

Bases: object

Parameters:
name: str
title: str
task: dict
agent: dict
protocol: dict
cells: tuple[Cell, ...]
validators: tuple[Validator, ...]
verdict: dict
bricks: dict
hypothesis: str | None = None
axes: dict
path: Path | None = None
property runs: int

Total executions this scenario asks for.

property declared_metrics: tuple[str, ...]
property manual_metrics: tuple[str, ...]

The metrics a human fills in, and the reason an output tree stays blind.

cell(name)[source]
Parameters:

name (str)

Return type:

Cell

declared(cell)[source]

Everything one cell declares, the values it falls back to included.

The single fallback rule, and the only one: runner.one_run builds a run from this and outputs.cell_fingerprint writes it into the ledger. Two implementations would let state.json promise one configuration while the agent received another, which is the drift this module spends its whole length refusing.

The bricks are resolved from their names, because [harness] is where a tag lives: a cell whose delta never changed still loads something else once the entry it names is repointed.

Parameters:

cell (Cell)

Return type:

dict

property reference: str

The cell every gap is measured against, as a cell name.

trysquare.scenario.load(path)[source]

Reads a scenario file and refuses it if it is not an experiment.

Parameters:

path (str | Path)

Return type:

Scenario

trysquare.scenario.parse(raw, path=None)[source]

Same as load, from an already-parsed mapping.

Split out so the rules can be tested without writing files.

Parameters:
Return type:

Scenario

trysquare.measure

What counts as a measurement, and how metrics are combined.

Two rules carry most of the weight here, and both exist because their absence produced published numbers that were wrong.

  1. A run counts only if it consumed tokens. A provider that cuts the stream leaves pi retrying and then returning turns that are real but empty. Such a run breaks no rule, touches no file and fails no test, so a naive harness records it as exemplary. Five of the bench’s defects were this same mistake in five disguises: “did not do the work” read as “worked well”.

  2. A validator that could not judge never yields a verdict. A crash, a timeout, unreadable JSON or a missing declared metric makes the run invalid, not false.

Pure: text and dictionaries in, dictionaries out. No subprocess, no file.

trysquare.measure.LINE_LIMIT = 8388608

The longest a line may be and still be read as an event. The agent’s stream is the one input here with no size anyone controls, and a reader that holds a whole line is only bounded if the writer emits newlines - which nothing promises.

class trysquare.measure.Run(id, cell, repetition, usage=<factory>, duration=0, metrics=<factory>, reasons=<factory>, state='valid', detail='', attempts=1)[source]

One execution of one cell, and everything known about it.

Parameters:
id: str
cell: str
repetition: int
usage: dict
duration: int = 0
metrics: dict
reasons: dict
state: str = 'valid'
detail: str = ''
attempts: int = 1
property is_valid: bool
property retries: int
trysquare.measure.decoded(lines)[source]

Every JSON object among these lines.

The one tolerance shared by every reader of these files. A line is decoded if it looks like JSON, whatever whitespace surrounds it, and anything else is skipped rather than fatal: a cut stream ends mid-line, and the lines before the cut are still evidence.

trysquare.measure.events(text)[source]

Every JSON object in a text held in memory.

For what is small enough to hold: an archived session, a fixture in a test. The agent’s own stream is not, and is read by read_file.

Parameters:

text (str)

trysquare.measure.one_line(text)[source]

Whitespace collapsed, so a detail stays on the line that reports it.

Every detail is printed as the tail of a one-line run report and stored in the ledger. A provider’s message arrives with the newline it was written with, which left a blank line under each failing run and a trailing n inside state.json.

Parameters:

text (str)

Return type:

str

trysquare.measure.counted(n, noun, plural=None)[source]

1 run, 2 runs, 0 runs, and 2 passes when s is not enough.

Here rather than beside the parser because four modules print counts, and the one that reached a real matrix - no price on 1 archived valid runs - was written where a helper living in the command line could not be reached.

A count of the form x of y takes its plural from y and does not come through here: 1 of 3 runs is already right.

Parameters:
Return type:

str

class trysquare.measure.Reading(usage, response, error, gave_up='')[source]

What one run’s stream is worth: the numbers, the answer, the first failure.

Everything anybody derived from a pi –mode json stream, and the reason the stream itself never has to be held: these are bounded where it is not. response is one assistant message, so its ceiling is the provider’s output limit rather than the length of the run.

gave_up is the error the run ended on, empty when it ended on anything else. pi closes a message it could not get from the provider with stopReason: “error”, and stops once its own retries are spent: a last message closed that way is a run the provider abandoned, however many turns came before it.

Parameters:
usage: dict
response: str
error: str
gave_up: str = ''
class trysquare.measure.Fold[source]

Everything a stream says, folded one event at a time.

The three rules are independent - a sum, a last, a first - so one forward pass answers what four separate walks over the stream used to answer. Stateful rather than a loop, because the same fold is fed two ways: read hands it a stream that has ended, and a live board hands it each event as it arrives. One implementation, so what a dashboard shows during a run and what the ledger records after it cannot be two different numbers.

Turns are counted on message_end events carrying a usage, not on turn_end. The usage filter is what makes a counted turn a billed turn: without it, a run that produced nothing still reports turns.

Retries matter beyond logging. When the stream is cut, pi replays the turn with the whole accumulated context, so input tokens, turns and duration inflate without the measured configuration having anything to do with it. Measured on the bench: zero retries gives 4 turns and 15.9k input tokens, thirteen retries gives 24 turns and 79.4k. Publishing those columns without looking at retries means publishing our own load on the provider.

feed(event)[source]
Parameters:

event (dict)

Return type:

None

reading()[source]
Return type:

Reading

trysquare.measure.read(evts)[source]

Everything a stream that has ended says.

Return type:

Reading

trysquare.measure.read_file(path)[source]

The same, over a trace on disk, holding one line at a time.

Read in binary and cut on b”n” alone, because text mode applies universal newlines and would cut on a bare r too - which a provider writes inside an error message, turning one decodable event into two fragments that decode as nothing. errors=”replace” for the same reason the session readers use it: one bad byte must not cost the whole file.

readline is bounded, and that bound is the point rather than a precaution. Nothing in the format promises a newline, and iterating a file whose writer emitted none rebuilds in memory exactly the object this reading exists to avoid. A line longer than the limit is not an event, whatever else it is, so it and what trails it up to the next newline are dropped.

An absent file reads as an empty stream: a spawn that failed wrote nothing, and that is silence rather than an error to raise here.

Parameters:

path (Path)

Return type:

Reading

trysquare.measure.bounded(handle)[source]

Every line of at most LINE_LIMIT bytes, longer ones dropped whole.

A chunk that neither ends in a newline nor stopped short of the limit is the head of a line too long to be an event; it and everything up to the next newline go.

Shared by the two readers of one stream - the sieve that writes a trace and the reading that measures it - so what counts as a line is decided once.

trysquare.measure.strip_session(session)[source]

Same numbers, from an archived session file instead of a live stream.

A session has a different shape from the stream: its line types are session, model_change, thinking_level_change and message, and usage is nested under message.usage. That shape difference is what is passed to _usage_sum, so comparing the two paths still tests the extraction - what it no longer tests is the addition, which is now written once.

retries is not recoverable here: it derives from auto_retry_start, a stream event that a session does not contain. Reported as None so a caller cannot mistake it for zero.

Parameters:

session (str)

Return type:

dict

trysquare.measure.thinking_levels(session)[source]

Every thinking level the session recorded.

Used by the smoke pass to check that the level a cell declared is the level that actually ran. The defect that made the thinking cell identical to the baseline in every published matrix cannot survive this check.

Parameters:

session (str)

Return type:

list[str]

trysquare.measure.models(session)[source]

Every model the session recorded, in order.

A scenario declares model as a pattern, which the agent resolves against the models the provider actually offers: gemma-4 ran as gemma-4-31b. So the declared value names an intention, and only the session names what answered.

The same reason etalon_commit is archived beside etalon: a name and what that name resolved to are two different facts, and an archive keeping only the name cannot say what it measured.

Parameters:

session (str)

Return type:

list[str]

trysquare.measure.strip(stream)[source]

The numbers alone, from a stream held in memory.

Parameters:

stream (str)

Return type:

dict

trysquare.measure.final_text(stream)[source]

The agent’s last piece of prose, from a stream held in memory.

Parameters:

stream (str)

Return type:

str

trysquare.measure.consumed_tokens(usage)[source]

The one thing that distinguishes a measurement from an incident.

Parameters:

usage (dict)

Return type:

bool

trysquare.measure.kind(value)[source]

How a metric can be aggregated, from its type alone.

Booleans become rates, numbers become medians, anything else is diagnostic: readable in a single run, never carrying a verdict. issues = [“#1”] says which issue overflowed, and there is no median of that.

Return type:

str

trysquare.measure.scorable(value)[source]
Return type:

bool

trysquare.measure.merge(results, declared)[source]

Combines what the validators returned into one measurement line.

results is [(validator mode, parsed JSON), …]. A validator that failed passes None as its payload.

Returns (metrics, reasons, state, detail). Every declared metric must be present; anything extra is kept but cannot be scored, which is what lets a general-purpose validator be reused across scenarios and lets a metric already paid for be scored later without remeasuring.

A validator may also name a metric under unjudged, meaning it could not judge that one while the rest of the run is fine. The name counts as returned but no value is recorded, so rate drops it from the denominator - which it always knew how to do, “out of how many could say” - and the reason is kept so the hole is readable. Recording false instead would file “could not judge” as “worked badly”, which is the one confusion this whole module is built against.

That the name is returned rather than simply omitted is what keeps the net tight: a typo produces a genuinely absent key, so it is still an invalid run.

Parameters:
Return type:

tuple[dict, dict, str, str]

trysquare.measure.fill_manual(run, manual)[source]

Applies hand-filled metrics, which may only fill a hole.

A form can supply a metric no mechanism produces. It can never overwrite a measured one: that is “the bench computes the verdict, the author does not work around it” applied to the interface. Returns the refusals.

Parameters:
Return type:

list[str]

trysquare.measure.valid_runs(runs, validity=())[source]

The runs a cell may be aggregated over.

validity names metrics that must be true for a run to count, declared by the scenario. delivered is the usual one: a run that modified nothing is not a disciplined agent, it is an agent that did not work, and without that filter it reads as a perfect score.

Parameters:
Return type:

list[Run]

trysquare.measure.rate(runs, metric)[source]

How many runs had the metric true, out of how many could say.

Parameters:
Return type:

tuple[int, int]

trysquare.measure.median(runs, key)[source]
Parameters:
Return type:

float | None

trysquare.measure.series(runs, metric)[source]

A cell’s per-run values for the criterion, as 0/1 or numbers.

This is what the verdict resamples. Booleans become 0/1 so one code path covers both a rate and a median.

Parameters:
Return type:

list[int]

trysquare.verdict

The executable half of the publication standard.

A gap reaches a page only if the harness certifies it. That is deliberate: the author cannot work around it, and six published conclusions collapsed on rerun while every one of them looked solid when it was written.

One mechanism covers both rates and medians: replay the draw. Resample the runs with replacement, recompute the gap to the reference cell on each draw, and the gap is publishable if its 95% interval excludes zero. Nothing else to read, no fragile statistic.

This judges a gap. An isolated measurement - the glow costs 23% of a frame budget - asserts no effect: it is published with its dispersion and no verdict.

The seed is fixed. A verdict that is not reproducible would make the harness itself a source of irreproducibility.

Pure: lists of numbers in, intervals out.

trysquare.verdict.mean(values)[source]

The statistic for a rate: booleans arrive as 0/1.

Parameters:

values (list[float])

Return type:

float

trysquare.verdict.gap_draws(reference, cell, stat=<function median>, draws=10000, seed=20260729)[source]

The resampled values of stat(cell) - stat(reference), sorted.

Draw order matters for byte-identical reproduction: the reference sample is drawn before the cell sample on every iteration.

Parameters:
Return type:

list[float]

trysquare.verdict.gap_interval(reference, cell, stat=<function median>, draws=10000, seed=20260729)[source]

95% interval of stat(cell) - stat(reference).

Parameters:
Return type:

tuple[float, float]

trysquare.verdict.interval(values, stat=<function median>, draws=10000, seed=20260729)[source]

95% interval of stat(values) itself, for a measurement that is not a gap.

The same mechanism, replaying the draw, so a dispersion is read exactly as an interval around a gap is. What it never gets is a state: an isolated measurement asserts no effect, so established would be a category error. A cost is published with its dispersion and no verdict, and the comparison that does carry one lives in the gap table.

Parameters:
Return type:

tuple[float, float]

trysquare.verdict.judge(reference, cell, stat=<function median>, draws=10000, seed=20260729)[source]

A gap, its interval, its p-value, and one of exactly two states.

Two states only. A third would invite a reading where a gap is “almost” something, and almost is how six conclusions got published. The p-value comes from the same draws as the interval and decides nothing. It exists so a whole table can be adjusted for the number of gaps it tests, which an interval cannot.

Parameters:
Return type:

dict

trysquare.verdict.holm(p)[source]

Holm’s step-down adjustment of p, in the order given.

Every gap of a table is its own test at 95%, so fifteen gaps with no real effect still yield 0.75 stars on average. Holm bounds the chance of even one false star over the whole table, and it assumes nothing about how the gaps are correlated. That matters here, since every column of a row resamples the same runs.

Parameters:

p (list[float])

Return type:

list[float]

trysquare.verdict.signed(x)[source]

A signed number with enough precision never to lie.

Rounding to integers displayed [-4, -0] for a bound worth -0.5: the reader believes they see an interval containing zero, so an inconclusive result, when the computation says the opposite. A non-zero bound must never render as zero.

Parameters:

x (float)

Return type:

str

trysquare.verdict.plain(x)[source]

An unsigned number, spaced like signed and as unwilling to round to zero.

A cost is a level rather than a difference, and a leading + on one would read as an increase over something the reader would then go looking for.

Parameters:

x (float)

Return type:

str

trysquare.verdict.probability(p)[source]

A p-value, never rendered as zero: resampling cannot prove a gap certain.

Parameters:

p (float)

Return type:

str

trysquare.verdict.points(x)[source]

A rate gap, in points. Rates live in 0..1 and read in 0..100.

Parameters:

x (float)

Return type:

str

trysquare.table

Rendering measures as a table, and gaps as verdicts.

Two tables, and they answer different questions. The score table says what each cell did, test by test. The gap table says which differences survive resampling, and it is the only one a sentence may rest on.

Both are computed from measures alone, which is what lets a rendering defect be fixed without paying for a matrix again.

Pure: runs in, strings out.

class trysquare.table.Measure(name, extract, stat, render)[source]

One column of the gap table.

Parameters:
name: str
extract: Callable[[Run], float]
stat: Callable[[list[float]], float]
render: Callable[[float], str]
trysquare.table.cost_measures()[source]

The columns that describe what a run cost.

Read these only alongside the retry count. When the provider cuts the stream, pi replays the turn with the whole accumulated context, so all four inflate without the measured configuration having anything to do with it.

Return type:

tuple[Measure, …]

trysquare.table.criterion_measure(criterion, sample=None)[source]

The column that carries the scenario’s criterion.

A boolean criterion is a rate and renders in points; a numeric one is a median and renders as a plain number. The type decides, because a scenario has no place declaring what Python already knows.

Parameters:
  • criterion (str)

  • sample (Run | None)

Return type:

Measure

trysquare.table.gap_rows(by_cell, reference, measures, validity=(), draws=None, seed=None)[source]

One row per cell, one verdict per measure, against the reference cell.

Only valid runs enter a verdict: a run that delivered nothing, or whose tests fail, measures neither its cost nor its discipline.

Each case also carries holm, its p-value adjusted over every case of the table, because that is the family a reader scans for stars.

Parameters:
Return type:

list[dict]

trysquare.table.gap_table(rows, reference, draws, seed)[source]

The gap table, in markdown.

Parameters:
Return type:

str

trysquare.table.retry_warning(by_cell)[source]

A note when the cost columns must not be read.

When the provider cuts the stream, the agent replays the turn with the whole accumulated context, so tokens, turns and duration inflate without the measured configuration having anything to do with it. Measured on the previous bench: zero retries gave 4 turns and 15.9k input tokens, thirteen retries gave 24 and 79.4k.

Publishing those columns without looking at retries means publishing our own load on the provider, so the table says so rather than relying on the reader to remember.

Parameters:

by_cell (dict[str, list[Run]])

Return type:

str

trysquare.table.spend_measures()[source]

The columns of the cost table: what a run cost, as a level rather than a gap.

Tokens in, tokens out and duration, which is what a reader asks for when deciding whether a configuration is affordable at all. turns is left to the gap table: it is a shape of the conversation rather than a price, and it is read against the reference or not at all.

The same caveat as cost_measures applies, and retry_warning states it: a retry replays the turn with the whole accumulated context, so all three inflate without the measured configuration having anything to do with it.

Return type:

tuple[Measure, …]

trysquare.table.spend_rows(by_cell, measures, validity=(), order=(), draws=10000, seed=20260729)[source]

One row per cell, one median and one 95% interval per measure.

Aggregated over the runs the verdict rests on - valid, and passing [verdict].validity - and not over every run that consumed tokens. A run that delivered nothing is cheap by construction, and averaging it into a price makes the configuration that fails most often look like the affordable one.

Each interval is computed from that cell alone. There is no reference here and therefore no state: a level asserts no effect, and the comparison that does carry a verdict is the gap table.

Parameters:
Return type:

list[dict]

trysquare.table.spend_table(rows, measures, draws, seed)[source]

The cost table, in markdown: cells in rows, median and interval in columns.

Parameters:
Return type:

str

trysquare.table.compare_rows(left, right, tests)[source]

One row per cell across two experiments, one x/n -> y/n per test.

Rates only, and no verdict. Costs are never tabulated across matrices: their runs did not see the same provider load, so the columns would compare our scheduling. And resampling certifies gaps within one matrix - a cross-experiment claim needs a scenario that measures both cells in one.

A cell present on one side only is named, with - on the other: dropping it would hide exactly the difference a comparison exists to show.

Parameters:
Return type:

list[dict]

trysquare.table.compare_table(rows, tests, left, right)[source]

The side-by-side score table, in markdown.

Parameters:
Return type:

str

trysquare.table.scored_metrics(runs, declared)[source]

Splits the declared metrics into those that count as a test, and the rest.

A test is a boolean: it passed or it did not, so a cell of the score matrix is x/n. A number has a median rather than a count, and a list has neither, so neither belongs in that matrix - but both are named rather than dropped, because a declared metric that vanishes from every table reads as a metric that was never measured.

Declaration order is kept: the scenario’s metrics contract fixes the column order, the same way [axes] fixes the row order.

Parameters:
Return type:

tuple[tuple[str, …], tuple[str, …]]

trysquare.table.score_rows(by_cell, tests, order=())[source]

One row per cell, one x/n per test.

Aggregated over the runs that produced a measurement, and not filtered by [verdict].validity: the validity metrics are columns of this very matrix, and a delivered column reading 10/10 by construction would hide the thing it is there to show.

n is per test rather than per cell, because a validator may return a metric as unjudged: it drops out of that one denominator and leaves the others intact. On a published matrix - complete, every run valid, nothing unjudged - n is the repetition count in every cell, and a n below it is exactly the signal to read.

Parameters:
Return type:

list[dict]

trysquare.table.score_table(rows, tests, other=())[source]

The score matrix, in markdown: cells in rows, tests in columns.

Parameters:
Return type:

str

trysquare.parity

Proving this harness reproduces the bench it replaces.

Parity is demonstrated in layers, and three of them are checkable exactly, at zero tokens, from material already archived. The question “what tolerance” only arose while we believed two samples had to be compared.

layer 1 stripping exact from archived sessions layer 2 scoring exact from tag + diff.patch layer 3 aggregation + verdict exact from the published per-run rows layer 4 launching the agent not comparable, it samples

Layer 3 needs nothing but a JSON file that already exists, so it could be verified before half of this tool was written. Layer 2 needs a tree, so it takes its reconstitution from the caller: the machinery belongs to replay, and the cost of a clone per run is the caller’s to declare.

Neither tool is the reference. Two computations are compared over the same archived material, and the material arbitrates. A gap on an exact layer blocks, and has exactly three admitted outcomes:

  1. this harness is wrong -> fix it

  2. the bench was wrong -> fix the published number and record the

    defect; the bench has twenty catalogued, so presuming it correct would be a losing bet

  3. archive artefact -> documented, removed from the parity scope

    (retries, which are a stream event a session does not contain)

class trysquare.parity.Report(observed=<factory>, problems=<factory>)[source]

What a layer found: what it observed, and what blocks.

The two are kept apart because “58/60 sessions reproduce exactly” is a layer narrating, not a layer failing. Folding both into one list of strings made a passing parity exit non-zero, and left the caller sniffing line endings to tell a count from a defect.

Parameters:
observed: list[str]
problems: list[str]
property holds: bool
property lines: list[str]
trysquare.parity.read_bench_measures(path)[source]

Loads the bench’s per-run rows, grouped by cell.

The published JSON is a list of per-run entries, not aggregates, which is the single fact that makes exact parity possible.

Parameters:

path (str | Path)

Return type:

dict[str, list[Run]]

trysquare.parity.published_by_id(path)[source]

The published rows keyed by run identifier, which is how an archive is laid out.

Rows without an identifier are dropped: they cannot be matched to anything on disk, so no layer that reads the archive can say a word about them.

Parameters:

path (str | Path)

Return type:

dict[str, dict]

trysquare.parity.archived_runs(rows, archive)[source]

The published runs an archive can reconstitute: those holding a diff.patch.

Exposed rather than inlined so a caller can say what a re-scoring is about to cost - one clone per run - without restating the rule that decides it.

Parameters:
Return type:

list[Path]

trysquare.parity.layer3(measures_path, reference='base', criterion='overflow')[source]

Recomputes the bench’s gap table from its own published rows.

Returns the rows, for a caller to compare against what the bench published. Identical inputs and an identical method must give identical output; anything else is one of the three outcomes above.

Parameters:
Return type:

list[dict]

trysquare.parity.layer1(measures_path, archive)[source]

Recomputes tokens and turns from the archived sessions.

A genuine test rather than a tautology: the bench counted message_end events in the live stream, and this reads the archived session, whose line types are session, model_change, thinking_level_change and message, with usage nested under message.usage. Two different paths to one number.

retries is deliberately not compared. It derives from auto_retry_start, a stream event a session does not contain, so it is an archive artefact and is removed from the parity scope rather than silently treated as zero.

Returns the differences, none when the layer holds.

Parameters:
Return type:

Report

trysquare.parity.layer2(measures_path, archive, reconstitute, validate)[source]

Re-scores archived runs by reconstituting their trees, and compares.

reconstitute(run_dir) -> Path rebuilds the tree from the tag and the archived diff and returns the context naming it; validate(context) -> dict returns the metrics. Both are injected so this stays testable and so the caller owns the cost of cloning.

Costs no tokens, which is what makes “fix a signature and re-score runs already paid for” true rather than aspirational.

A metric no validator returns for any run is a statement about scope, not a defect: the bench scored some metrics with a judge, whose verdict costs tokens and is not in this archive to be reused. Those are named once and left out, the way layer 1 leaves out retries. A metric returned for some runs and missing for others is the opposite - the validator is unreliable - and blocks.

Parameters:
Return type:

Report

trysquare.parity.layer4(experiment, workdir=None)[source]

The smoke pass: mechanical criteria that do not depend on the sample.

Layer 4 launches the agent, so its rates are a different sample and can never demonstrate parity. What it can demonstrate is that the harness is wired, and every criterion here is checkable without any statistical claim:

  • every run valid, meaning every run consumed tokens

  • the outputs a complete matrix owes are present

  • each run’s directory holds its context, configuration, diff and validation

  • the thinking level each session recorded equals the level its cell declared - the check that makes the defect which rendered the thinking cell identical to the baseline unable to recur

  • the model that answered is the one the scenario’s pattern asked for, which is the same check one axis over: model is a pattern, and a pattern that quietly resolved elsewhere measures a model nobody declared

Returns the failures, none when the pass holds.

Parameters:
Return type:

Report

trysquare.parity.compare(rows, expected)[source]

Checks recomputed rows against the values the bench published.

Compares the rendered strings rather than raw floats: what was published is what a reader saw, and a parity that agrees on invisible digits while disagreeing on the printed ones would be worthless. Returns the differences, empty when parity holds.

Parameters:
Return type:

list[str]

The effectful shell

trysquare.config

The config file, and the hard rule about what it may not contain.

There are no environment variables anywhere in this tool. Its predecessor had ten, and an environment variable is invisible inheritance: the reader of a scenario cannot see it, the archive does not record it, and the value that actually ran is whatever the shell happened to hold.

A config file may only supply machine paths and load fallbacks. Provider, model, thinking level, etalon and repetitions are mandatory in the scenario and raise when absent. If they could be inherited from here, the same scenario file would measure something different on another machine, and that is precisely the defect that made the thinking cell identical to the baseline in every published matrix.

A repository entry may be a directory or a git URL. Both are addresses, and an address decides where the code is read from, never what is measured of it.

Precedence: scenario (the experiment) > CLI (explicit and announced) > config (the machine) > built-in defaults.

exception trysquare.config.ConfigError[source]
trysquare.config.closest(name, known)[source]

A parenthetical to append to a refusal, when the name looks like a typo.

difflib’s default cutoff keeps a far miss silent: suggesting rule for banana would decorate every honest refusal with noise, and a suggestion that is usually wrong teaches the reader to skip all of them.

Parameters:

name (str)

Return type:

str

trysquare.config.which_file(raw)[source]

Tells the two TOML files apart, or says it cannot.

Returns “config”, “scenario”, or None when the file is too empty or too mixed to say. Silence is the useful part: a scenario legitimately carries [harness] to pin bricks by tag, so a file with sections of both kinds is a scenario with an ordinary mistake in it, and the caller’s own refusals name that mistake better than a guess about which file it is.

Parameters:

raw (dict)

Return type:

str | None

class trysquare.config.Config(repos: 'dict' = <factory>, harness: 'dict' = <factory>, defaults: 'dict' = <factory>, path: 'Path | None' = None)[source]
Parameters:
repos: dict
harness: dict
defaults: dict
path: Path | None = None
repo(name)[source]

Resolves a logical repository name to a directory on this machine.

A scenario writes repo = “my-repo”. It carries no machine path, which is what makes it portable and keeps one author’s directory layout out of an experiment file.

Raises when the entry is a URL: a URL has no local directory until it has been cloned, and answering with a plausible path that nothing has created yet is how a caller ends up handing git something that is not there.

Parameters:

name (str)

Return type:

Path

harness_repo(name)[source]
Parameters:

name (str)

Return type:

Path

remote(name)[source]

The git URL a logical name points at, or None when it names a directory.

The URL is returned verbatim. It is not expanded: a username or a token taken from the shell is invisible inheritance in its purest form - it does not appear in the archive, and the value that actually ran is whatever the environment happened to hold. Abolishing that is what this module is for.

Parameters:

name (str)

Return type:

str | None

harness_remote(name)[source]
Parameters:

name (str)

Return type:

str | None

workdir()[source]
Return type:

Path

fallback(key)[source]
Parameters:

key (str)

trysquare.config.is_remote(value)[source]

Whether a repository entry names a git URL rather than a directory.

This exists so a URL never reaches expand(). Path(“https://host/x”) collapses the double slash into https:/host/x, which is a relative path: git would then be handed a directory name resolved against the config file’s parent, and the failure is a clone of nothing rather than an error anyone can read.

file:// counts as remote. Path() mangles it exactly the same way, and treating it as a URL is also what lets the pinning path be tested end to end without a network.

Parameters:

value (str)

Return type:

bool

trysquare.config.expand(value, relative_to=None)[source]

Expands ~, $TMPDIR and friends, then anchors relative paths.

A relative path in a config file is relative to that file, not to the current working directory: the config describes a machine, and where the operator happens to stand when running a command is not part of it.

Only ever called on a path. A URL is kept verbatim - see is_remote.

Parameters:
  • value (str)

  • relative_to (Path | None)

Return type:

Path

trysquare.config.load(path=None, start=None)[source]

Reads the config file, or returns built-in defaults when there is none.

An absent config file is not an error: a scenario that names no logical repository needs nothing resolved.

Parameters:
Return type:

Config

trysquare.config.discover(start)[source]

Walks up from start looking for a config file.

Walking up is a convenience for the operator, never a way to inherit a measurement: what is found here can only ever be a path or a load fallback.

Parameters:

start (Path)

Return type:

Path | None

trysquare.repo

Preparing the repository a run measures, and injecting the harness bricks.

Two decisions here are not interchangeable with the obvious alternatives, and both come from a trap that was actually hit.

Clone at an etalon, never copy the working tree. The measured state has to be immutable and named, or yesterday’s measures no longer compare with tomorrow’s. And a repository may be a worktree, whose .git is only a file pointing at a shared gitdir: a recursive copy then gives every run the same gitdir, so one agent running git commit moves the comparison base of every run in flight and nobody notices.

A tag is named, not immutable. git tag -f moves one, and a matrix measured before the move no longer compares with one measured after - the two report the same etalon and measured different code. That is not hypothetical: two campaigns of the same scenario were joined by name and turned out to have read two different versions of the ticket under test, which commit_of was the only thing to record. So an etalon may also be a commit, written out in full, and is_commit decides which of the two it is. A commit cannot move, and the directory name then says what was measured rather than what it was called.

A remote is pinned as a working tree, never as a bare mirror. A bare mirror is smaller and answers git show and git ls-tree perfectly well, so it looks right. But etalon.checkout is walked as files by a validator reading the reference state, a bare repository passes an is_dir() check and yields an empty reference, and the whole matrix then reports plausible numbers about nothing. See pin.

Exclude what we injected. The files the harness drops are not the agent’s work. Without .git/info/exclude, scope scoring counts them as changes the agent made, and every configured cell drops to zero: a measurement of our own tooling rather than of the behaviour under test.

Commit what we gave the task. A files brick is the one exception to the line above, and the difference is who the material is for. Harness plumbing is addressed to the agent library and belongs outside the repository’s history; a file a brick puts in the tree - a probe, a fixture - is addressed to the task, and the agent may edit it or delete it. Excluded, that tampering would leave no trace anywhere: git ignores an untracked path whether it was changed, weakened or removed. Committed, the given state becomes the base the diff is taken against, so the injection still costs nothing in touched and every later move on it is recorded. It is also what makes a replay exact - see inject.

exception trysquare.repo.RepoError(*args, detail='')[source]
Parameters:

detail (str)

class trysquare.repo.Prepared(path, etalon, injected=<factory>, given=<factory>, agents=<factory>)[source]

A clone ready to be measured, and what we put in it.

Parameters:
path: Path
etalon: str
injected: list[str]
given: list[str]
agents: dict[str, dict]
trysquare.repo.git(args, cwd=None, check=True)[source]
Parameters:
Return type:

str

trysquare.repo.is_commit(etalon)[source]

Whether this etalon names a commit outright rather than a ref.

Parameters:

etalon (str)

Return type:

bool

trysquare.repo.clone_argv(source, etalon, target, keep_tags=False)[source]

The flags a clone is made with, separated so they can be asserted.

–single-branch –branch <etalon> is the reproducibility guarantee: the clone is the pinned state and nothing else.

–no-tags is right for a run’s clone, where nothing needs the tag ref once HEAD is detached on it. A pinned source keeps its tags, because every run clones from that directory by tag name.

Measured rather than assumed: git does in fact keep the tag named by –branch even under –no-tags, so a pinned source would probably work either way. Nothing documents that interaction, though, and the pinned source’s whole job is to answer a clone by tag - so it does not rest on undocumented behaviour. The cost is the tag refs of one branch.

A commit takes none of that. –branch refuses anything that is not a ref, so a commit etalon is fetched whole and reached by the checkout clone runs next. That also settles the two flags: –single-branch would keep only the default branch, and –no-tags only the branch refs, either of which can leave the wanted commit unreachable in a clone that otherwise looks complete. Fetching everything is the price of an etalon that cannot move.

Parameters:
Return type:

list[str]

trysquare.repo.clone(source, etalon, target, keep_tags=False)[source]

Clones source at etalon into target, tag or commit.

source may be a local directory or a git URL. A URL is handed to git verbatim: resolve() would turn it into a path, which is the defect config.is_remote exists to prevent.

The exists() check stays as a backstop. runner.prepare_source checks earlier and says more, but a caller that reaches here with nothing on disk should still be told, not left with a bare git error.

A commit etalon leaves HEAD detached on it, which is where a tag etalon already leaves it: everything downstream - commit_of, etalon_file, etalon_files, the diff a run is scored on - reads the etalon as a revision and cannot tell the two apart.

Parameters:
Return type:

Path

trysquare.repo.pin(url, etalon, target)[source]

Clones a remote at etalon once, as a working tree runs can clone from.

A working tree and not a bare mirror. etalon.checkout is walked as files by a validator reading the reference state - which is what Assay.sources_at_etalon does - and a bare repository passes an is_dir() check, yields an empty reference, and turns every comparison against the etalon into a plausible number about nothing. A bare mirror would have been smaller and would have measured nothing, silently.

Parameters:
Return type:

Path

trysquare.repo.commit_of(source, etalon)[source]

The commit an etalon designates, for the archive.

Peeled with ^{commit} so an annotated and a lightweight tag record the same kind of object: two archives naming different object kinds for the same tag are not comparable. Also the only trace left when a tag is moved between two matrices - the reason an etalon may now be a commit outright, see is_commit.

Parameters:
Return type:

str | None

trysquare.repo.etalon_file(source, etalon, path)[source]

One file’s content at the pinned tag, without checking anything out.

Scoring needs the reference side of a comparison, and reading it from the tag means the reference cannot drift while a matrix is in flight.

Parameters:
Return type:

str

trysquare.repo.etalon_files(source, etalon, pattern='')[source]
Parameters:
Return type:

list[str]

trysquare.repo.inject(prepared, context=None, system=None, agents=None, skills=None, files=None, agent_model=None)[source]

Drops the harness bricks into the clone and records what was dropped.

files is the only one that lands in the measured tree, and it is committed rather than excluded. See the module docstring for why, and give for what the commit buys a replay.

Parameters:
Return type:

Prepared

trysquare.repo.give(prepared, files)[source]

Puts a brick’s files in the measured tree, and commits them at the etalon.

The commit is what separates given to the task from written by the agent without hiding anything. changed_files and diff both compare against HEAD, so a file committed here costs nothing in scope scoring - and the moment the agent edits or deletes it, that shows up in the patch like any other change. A probe handed to an agent is exactly the material an agent may be tempted to weaken until it passes, so it is the last thing that should be invisible.

It is also what makes the run replayable. A reconstitution clones the tag and applies the archived patch; the patch’s context lines for a given file only match if the same file is put back first, and committed the same way. cli.reconstitute calls this before apply_diff for that reason.

Never replaces what the tag holds. Overwriting a tracked file would change the measured code while looking like nothing at all: the diff is taken against a HEAD that already contains the replacement, so the substitution would be invisible in the archive and every cell would be measured on a repository nobody described.

Parameters:
Return type:

Prepared

trysquare.repo.agent_frontmatter(path, override=None)[source]

Reads an agent definition’s frontmatter, and settles which model it runs.

A subagent that declares no model inherits the operator’s defaultModel. Nine shipped agents were in that position, and a session on one provider ran them all on another and returned 402s. So the model is resolved here, and where it came from is recorded: the declaration may live in two places, but the trace settles which one applied.

Parameters:
  • path (Path)

  • override (str | None)

Return type:

dict

trysquare.repo.check_agent_models(agents)[source]

Refuses to measure a subagent that would run on an inherited model.

Parameters:

agents (dict[str, dict])

Return type:

None

trysquare.repo.changed_files(d)[source]

Files the agent touched: new files and deletions included.

–intent-to-add is what makes new files appear at all: without it, a change written into a file created for the occasion is invisible. It writes to the index, which is safe because every run owns its own clone.

Parameters:

d (Path)

Return type:

list[str]

trysquare.repo.diff(d)[source]

What the agent changed, in a form that can be applied again.

–binary is what makes it applicable. Without it git records a binary file as Binary files /dev/null and b/x.pyc differ, which carries no content and an abbreviated index, and git apply refuses it - and refuses the whole patch, the source changes with it. An agent that runs the declared suite to check its own fix leaves __pycache__/*.pyc behind, so this is the common case, not the exotic one: the archive looks fine until a replay –rescore months later cannot use it.

The extra bytes are base85 of what the agent produced, and only for the files it produced. A diff nobody can apply is not smaller, it is empty.

Parameters:

d (Path)

Return type:

str

trysquare.repo.apply_diff(d, patch, what='')[source]

Replays an archived diff onto a fresh clone.

This is what makes a validation replayable months later: the archive keeps the tag and the patch, not 150 copies of a working tree.

what names whose patch it is. A replay walks every run in a directory, so a refusal that says only “the archived diff” leaves the reader to find which of sixty it was.

Parameters:
Return type:

None

trysquare.agent

Building the agent invocation, and running it.

The argument list is the experiment. Everything here is explicit on purpose: discovery is switched off wholesale and every brick is then handed back by path. Discovery is gated on project trust, it walks up ancestor directories, and it fails silently - three ways for a cell to measure the absence of the brick it believes it is measuring.

The stripping of the output stream lives in measure.py, because it is a pure function from text to numbers and belongs where it can be tested without a subprocess.

trysquare.agent.resolves_to(declared, ran)[source]

Whether the model that ran is the one the scenario’s pattern asked for.

–model takes a pattern, not an id: a scenario declaring gemma-4 ran as gemma-4-31b, which is resolution and not substitution. So an equality check would refuse every legitimate run, and no check at all would let a fallback to the machine’s defaultModel pass unseen. What is verified is the weaker, checkable property: the declared pattern must still be in what answered.

Both sides may carry a provider/ prefix, and a declared pattern may carry the :<thinking> shorthand the agent also accepts; neither says anything about which model ran, so both are stripped before comparing.

Parameters:
Return type:

bool

class trysquare.agent.Outcome(trace, response, error, stderr, code, duration, timed_out, usage, overflowed=False, gave_up='')[source]

One invocation of the agent, whatever happened to it.

The stream itself is not here, and that absence is the point. It is the one thing the harness handles whose size nobody controls, and holding it is how a single runaway run reached 136 GB and had the matrix killed under it. What is kept is everything that was ever derived from it - the numbers, the answer, the first failure - each bounded by one message rather than by the length of the run. The stream stays where it was written, and trace says where.

Parameters:
trace: Path
response: str
error: str
stderr: str
code: int | None
duration: int
timed_out: bool
usage: dict
overflowed: bool = False
gave_up: str = ''
property produced_something: bool

Whether this is a measurement at all.

A run stopped for writing past its ceiling is not one, even when the bytes it managed before derailing carry a usage. Recording it would publish the cost of a run that never finished - the founding confusion of this harness, “did not do the work” filed as “worked well”, reached by a new road.

A run the provider gave up on is not one either, for the same reason. The turns before the failure are real, and they are exactly what made three whole cells of a matrix read as agents that chose to change nothing.

property signalled: bool

Whether something outside the run ended it.

A negative status is the signal that killed the child. It is not a result the agent produced, and it is not the agent’s silence either - which matters because those two are told apart by the same emptiness. The OOM killer is the case that shows the difference: it used to buy three agent runs in a row, each one killed for the same reason as the last.

trysquare.agent.argv(prompt, provider, model, thinking, session_dir, extensions=None, skills=None, has_context=False)[source]

The exact argument list for one run.

  • -a is unconditional. In non-interactive mode there is no trust prompt, so without a saved decision every .pi/ resource in a fresh clone is ignored silently. A fresh clone never has a saved decision. Passing it is a validity condition, not a comfort.

  • -ns -np -ne switch off skill, prompt-template and extension discovery; explicit –skill and -e paths still load, and they are the only way a brick enters.

  • -nc unless the cell provides a context file. There is no –context-file in the agent, so an AGENTS.md can only be obtained by writing it into the clone and letting discovery run. Discovery walks up ancestors, so cells without context are protected from inheritance and cells with context are not. That asymmetry is known and declared rather than hidden.

  • –thinking always, never inherited.

Parameters:
Return type:

list[str]

trysquare.agent.DISCARDED = b'{"type":"message_update"'

The one event kind a trace does not keep. It carries the whole message accumulated so far, twice - once as partial and once as message - to deliver a delta of two characters, so a stream costs the square of what the agent says. Nothing reads it: every measurement comes off message_end, and the validator that once parsed these was reworked to read the session because parsing them was the defect.

Matched as a prefix rather than searched for, so a message_end whose text happens to quote this name cannot be dropped. Should pi ever reorder its keys the match stops, the trace grows back, and the ceiling says so - the failure is a loud one.

trysquare.agent.run(cwd, args, timeout, trace, ceiling=None, watch=None)[source]

Runs the agent once through the sieve, and reads back what it says.

Nothing here ever holds the stream: it passes line by line and only what an event is lands in trace. That matters because the stream is quadratic in the length of the answer - pi sends the whole accumulated message twice on every update, to deliver two characters - so a 126 KB reply costs a gigabyte and a 1.5 MB one costs the 136 GB that started all this.

ceiling bounds what is kept, which is the honest quantity now. Bounding the raw stream would keep excluding the cells whose agent answers at length, and since the cost goes as the square, a ceiling with twice the bytes leaves only half again as much to say.

stdin must be closed. With an open pipe, the agent waits indefinitely for something to read and the run freezes without emitting a byte.

Parameters:
Return type:

Outcome

trysquare.agent.run_until_productive(cwd, args, timeout, attempts, trace, ceiling=None, watch=None)[source]

Retries only while nothing has been produced.

A run that consumed no tokens produced no result, so there is nothing to select between: retrying it is not optional stopping. A run that did produce something is never retried, whatever its result.

A run something else ended is not retried either. It looks identical to silence from here - no tokens, nothing to select between - and the loop used to answer it by launching a fresh agent, which is how one Ctrl-C bought three more.

A run stopped for overrunning its ceiling is not retried either, and for the same reason as a signalled one: the next attempt would be the same runaway, and three of them is the incident this ceiling exists to end rather than to triple.

Every attempt writes the same trace, so it holds the attempt that was kept. The sessions are what say what the others did: they are archived one file per attempt.

Parameters:
Return type:

tuple[Outcome, int]

trysquare.agent.export_html(session, target, timeout=120)[source]

Renders one archived session as a standalone page, by the agent itself.

The agent already knows how to read its own sessions, so nothing here reimplements that: a renderer written here would drift from the format it renders, silently, and the format is the agent’s rather than ours.

pi –export takes no output path and writes pi-session-<stem>.html into the current directory, so the target directory is the working directory. The file is then renamed to <stem>.html, which puts the page beside the jsonl it came from under the same stem - the archive stays readable by looking at it.

–offline because an export reads a file. A startup network call would make re-rendering an archive depend on the network being up, which is the opposite of what an archive is for.

Raises RuntimeError on anything that went wrong, so a caller has one exception to catch and one session’s failure need not cost the others.

Parameters:
Return type:

Path

trysquare.agent.ambient_thinking(settings=None)[source]

The thinking level a subagent will actually run at.

A subagent’s level cannot be declared: the frontmatter has no field for it and the library passes no option, so it always comes from the operator’s settings. Read here so the harness can check a scenario against it and refuse rather than produce a matrix whose cells claim one level and ran another.

Parameters:

settings (Path | None)

Return type:

str | None

trysquare.agent.available()[source]

Whether the agent binary is on PATH, for a clear message rather than a trace.

Return type:

bool

trysquare.validation

Running the validators, and keeping them honest.

Validators are independent: each receives the same context and cannot see what the others found. That is not tidiness. A judge told the script’s verdict is anchored on it, and its agreement stops being an independent signal - which was the only reason to have a judge at all.

A judge is also blind: its context carries no cell name and no configuration. Blinding is not always achievable, though, and pretending otherwise would be worse than saying so: when the treatment is the prompt, handing the judge the prompt reveals the cell. So the harness reports which declared pieces vary across cells, and the operator knows how blind the judge really is.

class trysquare.validation.Result(mode, payload, stderr='', detail='')[source]

What one validator returned, or why it did not.

Parameters:
mode: str
payload: dict | None
stderr: str = ''
detail: str = ''
property ok: bool
trysquare.validation.where(path)[source]

A path as a context carries it: absolute, always.

“Every path it needs is absolute in the context” is a documented promise, and it is what lets the child run somewhere that is deliberately not the measured clone. It was true by accident rather than by construction: a run’s paths come from the work directory, which the config makes absolute, so nothing relative had ever been passed.

replay passed one. Its archive directory is whatever the operator typed - results/… - so the archived session went into the context relative, the validator’s child resolved it from its own working directory, found nothing, and reported that the run had no session. Which reads as a fact about the agent: sixty runs said “nothing about the agent’s process can be read” and the metric of process this file’s docstring calls replayable was unjudged on every one of them.

Resolved here, once, so no caller can be the one that forgets.

Parameters:

path (Path | str)

Return type:

str

trysquare.validation.write_context(directory, repo, etalon, etalon_checkout, prompt_file, session_dir, trace, cell, repetition, blind=False, response_file=None, test_command=None, prepare=None, artefacts=None, touched=None, files=None, given=None, declared=())[source]

Writes the context file a validator is handed.

response is the agent’s final prose, extracted once by the harness. A validator that needed it would otherwise have to reimplement stream parsing, and every validator reimplementing it is every validator getting it slightly differently.

test_command is the suite the scenario declared, carried here for that same reason and one more: a validator that guessed it would be reading package.json, a file inside the perimeter the measured agent may edit.

Carried as the scenario wrote it - a string - rather than pre-split. One fact, one representation: an archived context read against the scenario file six months later says the same thing, with no transformation to know about. And a validator gets the shape its own runtime prefers, which is the opposite of what pre-splitting assumed: a shell splits a string for free, where a JSON array has to be parsed and rebuilt. Only Python prefers an argv, and a Python validator never sees this key - run.tests() does the splitting, with scenario.split_command, which is the same rule the loader vetted it with.

It is not withheld from a blind context. It is a property of the task, identical in every cell, so it tells a judge nothing about which configuration produced the work in front of it.

Absent when the scenario names no suite, and an absent key is a different fact from an empty command: it says this experiment scores no test suite, which is something a validator may need to refuse over rather than score.

artefacts names what running that task leaves behind and is not the agent’s work - declared for the same reason the suite is, since only the task’s author knows which paths in this repository are by-products. It filters what a verdict rests on and never what is recorded: touched stays complete, because hiding a measurement is the other dishonesty this harness refuses.

touched and files are the two facts every validator wanted and each computed for itself. repo.changed_files, repo.etalon_files and repo.etalon_file had held that knowledge all along, and three shipped validators reimplemented it with a raw subprocess regardless - one of them landing on a different answer for the reference side. Computed here once, they cannot be got slightly differently by three callers.

declared is the metric names the scenario contracted for. The harness already refuses a run whose validator omitted one, but only after the tokens are spent; handing them over lets the base say which one is missing before anything is recorded. Safe for a blind context: metric names say nothing about a cell.

Parameters:
Return type:

Path

trysquare.validation.run_script(validator, context, timeout, cwd=None)[source]

Runs a script validator: one argument, JSON on stdout.

The working directory is deliberately not the measured clone: a validator that wrote a stray file there would be counted as the agent’s work by scope scoring. Every path it needs is absolute in the context file.

Including the context file’s own path, which is the whole reason it is resolved here. This call changes the child’s working directory, so a relative path handed to it is measured from somewhere the caller never named - and –output out is the documented way to invoke the tool, so every script validator failed with unreadable context for want of one resolve().

Parameters:
Return type:

Result

trysquare.validation.judge_dossier(directory, validator, rubric, pieces)[source]

Writes what the judge is given, and returns its working directory and prompt.

The pieces are declared in the scenario, because what the judge is given to read is half of what it measures. Nothing about the cell is included: the judge must not know which configuration produced the work it scores.

Parameters:
Return type:

tuple[Path, str]

trysquare.validation.run_judge(validator, directory, prompt, brick, timeout, trace, attempts=1, ceiling=None)[source]

Runs the judge, and reads back the verdict its tool call recorded.

Retried only while there is no usable answer, which is not optional stopping: an absent verdict is not a verdict one could have preferred. Once a verdict exists it is kept, whatever it says.

trace is where the judge’s own stream goes, and it is given rather than derived from directory: directory sits inside the published archive, which deliberately never carries a raw stream. Nor is it discarded - the judge’s first error is the only diagnosis a silent judge leaves behind.

Parameters:
Return type:

Result

trysquare.validation.blindness(scenario)[source]

How blind a judge actually is in this scenario, piece by piece.

A piece that varies between cells leaks the treatment. Reported rather than forbidden: a judge on a matrix of prompts stays possible, provided it is said.

Parameters:

scenario (Scenario)

Return type:

dict

trysquare.validation.describe_blindness(report, cells)[source]

The lines printed at launch, so the operator is never surprised.

Parameters:
Return type:

list[str]

trysquare.validation.check_thinking_precondition(declared, ambient, uses_subagents)[source]

Refuses a scenario whose subagents would think at another level.

A subagent’s thinking level cannot be declared anywhere: the frontmatter has no field for it and the library passes no option. What cannot be controlled is verified instead, and a mismatch stops the run rather than producing cells that claim one level and ran another.

Returns the refusal message, or None when there is nothing to refuse.

Parameters:
  • declared (str | None)

  • ambient (str | None)

  • uses_subagents (bool)

Return type:

str | None

trysquare.assay

The one module a validator author imports. Everything else on this page is the harness; this is what the thing being run is written with.

The base a validator is written with.

An assay is an analysis performed on a sample, which is exactly the trade: one finished run goes in, named metrics come out. The three obvious names were taken - validation is the harness running validators, measure and verdict own the aggregation - and the distinction is worth keeping: that module is the caller, this one is what the callee is written with.

Four names, and an author needs no others:

from trysquare.assay import validator, Assay, Metric, CannotJudge

@validator
def evaluate(run: Assay) -> dict:
    outside = run.touched - {"counter.py"}
    return {
        "delivered": bool(run.touched),
        "in_scope": Metric(not outside, f"also touched {', '.join(outside)}"),
        "tests": run.tests(),
    }

if __name__ == "__main__":
    raise SystemExit(evaluate.cli())

An attribute costs nothing, a parenthesis costs something. What the harness already computed is an attribute (run.touched); what has to go and do work is a method (run.tests()). The frontier between the two therefore does not have to be remembered, it is read - and the harness pre-computes exactly what it can compute once for every language, because a fact computed in one place cannot drift.

Three states, and each has one way to say it. Getting this wrong is the mistake the whole project is built against, so the type makes the confusion inexpressible: a Metric has no value meaning “I could not tell”.

I judged it

a value, or Metric(value, reason)

not this metric

Metric.unjudged(why)

not this run

raise CannotJudge(why)

The last two both end as a non-zero exit and an invalid run, because from the harness’s side there is only one failure mode. The distinction is diagnostic, for whoever reads validation/<mode>.stderr six months later, and it lives in the wording rather than in the control flow.

trysquare.assay.summarise(output)[source]

The runner’s own summary, or a tail that says it is a fallback.

Anchored on the marker and taken to the end, because every one of the four runners writes its summary last. Nothing is parsed.

A fallback that does not declare itself a fallback is what let a real regression live: the shipped validator grepped for not ok, Node v23 changed its default non-TTY reporter from tap to spec, the grep stopped matching, and a silent tail[-1] returned a closing brace as the reason for a failing suite. Nothing looked broken for two major versions.

Parameters:

output (str)

Return type:

str

exception trysquare.assay.CannotJudge[source]

This run cannot be scored, which is not the same as scoring it badly.

Raised by a validator, and by the base whenever it is asked for something the context does not carry. It exits non-zero with a sentence and no traceback: it is not a defect, so a trace would only invite reading it as one.

exception trysquare.assay.ProbeTimeout[source]

A probe that ran too long, which the base declines to interpret.

A CannotJudge by default, because refusing is the safe reading. But a probe is usually milliseconds, so exceeding a generous timeout often means the agent’s fix loops forever - which is a failure of the work, not an inability to judge. Only the domain knows which, so it can catch this and score a failure with a reason that says so. The base does not decide on its behalf, and the shipped validator that hardcoded the choice could not express the other one.

class trysquare.assay.Metric(value=None, reason='', judged=True)[source]

A value and, when it helps, the reason it came out that way.

Two facts forced this shape rather than a subclass of int or bool. bool is final in Python, so a value that is both a boolean and carries a reason cannot exist - and the boolean is the common case. And metrics cross a JSON boundary, so measure.kind only ever sees a dehydrated value: an envelope cannot break the aggregation because it never reaches it. The entry point unwraps before writing.

A reason is published whether the value reads as a success or a failure. Filtering on failure is not implementable honestly: “failed” is only definable for a boolean, and cited_paths = 7 is neither a success nor a failure - its verdict comes from the gap table, not from here. A base that filtered would be wrong about every median.

Parameters:
value: object = None
reason: str = ''
judged: bool = True
classmethod unjudged(reason)[source]

This one metric could not be judged; the rest of the run is fine.

The name is still returned, which is what keeps the harness’s net tight: a typo in a metric name still produces a genuinely absent key and therefore an invalid run, loudly, while an honest “I cannot say” shrinks a denominator instead - visibly, since a rate renders as 7/8.

The case this exists for: a probe that could not run. issue1.py:310-316 returns {“ok”: False, “erreur”: “no game/ in the clone”}, which the harness records as par_face = false. That is “could not judge” filed as “worked badly”, on the metric carrying the verdict.

Parameters:

reason (str)

Return type:

Metric

trysquare.assay.plain(value)[source]

A metric value reduced to what JSON should carry.

A set becomes a sorted list, and the sort is a requirement rather than a tidiness. PYTHONHASHSEED is random by default, so an unsorted set of strings serialises in a different order from one process to the next: two identical measurements would produce byte-different measures.json, git diff would show churn that means nothing, and compare and parity would read a difference that is not there.

trysquare.assay.report(returned)[source]

Splits what a validator returned into the payload the harness reads.

unjudged is a third key beside metrics and reasons, and it is an addition to the contract rather than a change: a validator that never uses it produces exactly what it produced before. The harness moves those names out of the aggregation while still counting them as returned.

Parameters:

returned (dict)

Return type:

dict

class trysquare.assay.ToolCall(name, arguments, failed)[source]

One call the agent made, as the archived session recorded it.

Parameters:
name: str
arguments: dict
failed: bool
wrote(path)[source]

Did this call write to path?

A failed call never counts. pi rejected two edit calls carrying no path on a real run, and counting those would date the work before it happened.

The tool vocabulary ages with the agent and not with trysquare, so it ages loudly: a name this does not classify raises rather than answering “no”. A new writing tool would otherwise make a process metric quietly false, and a column that dropped would read as a less disciplined agent - a wrong conclusion rather than a visible hole.

Parameters:

path (str)

Return type:

bool

class trysquare.assay.Assay(context)[source]

One finished run, and everything a validator may ask about it.

Every attribute and every method goes through _part, which is the single seam the fake replaces. That is the whole reason for the indirection: a fake that overrode one accessor at a time would drift from the real one silently.

Parameters:

context (dict)

property repo: Path

The clone that was measured, with the agent’s work in it.

property etalon: str

The tag the clone started from.

property touched: frozenset[str]

The files the agent changed, new files included.

Computed by the harness, once, with git add -A –intent-to-add followed by git diff –name-only. Two shipped validators each reimplemented it and each copied the same comment; repo.diff had held the same knowledge all along.

property artefacts: frozenset[str]

The files it touched that the scenario declared as by-products of the task.

A subset of touched, never a replacement for it: what a verdict rests on is filtered, what is recorded is not. Subtract it to get the work - run.touched - run.artefacts - which is what a scope or a delivery metric wants.

Empty when the scenario declared nothing, so a validator written against this keeps working on a scenario that has no by-products to name.

Why declared and not guessed: only the task’s author knows which paths in their repository are by-products, and a built-in list would eventually hide a file an agent really wrote. Measured on a real matrix - an agent ran the declared suite to check its own fix, __pycache__/*.pyc counted as its work, and in_scope was false in every run of every cell. Not at random either: the runs that scored badly were the ones where the agent verified itself.

property given: frozenset[str]

The paths a files brick put in the tree before the agent started.

Empty for every cell that was handed nothing, which is most of them, and that emptiness is the fact a validator needs: a probe that is not there was either never given or deleted along the way, and those two are not the same measurement. Without this, a run that removed the test it was handed would score exactly like a run that was handed no test at all.

A path in here is tracked in the clone, committed on top of the etalon, so any edit the agent makes to it appears in touched like any other change.

A part rather than a plain read of the context, so a fake has to declare it: Assay.fake() answering “nothing was given” to a validator that never said it cared would put the absence of a fact into the shape of a fact.

property files_at_etalon: tuple[str, ...]

Every path in the pinned tree, unfiltered.

The list is cheap - one git ls-tree - and depends on nothing but the run, so the harness computes it. Reading their contents depends on a filter that belongs to the question, so that stayed a method.

property response: str

The agent’s final prose, extracted once by the harness.

A validator that parsed the stream itself would be one more validator getting it slightly differently.

property prompt: str
property declared: tuple[str, ...]

The metric names the scenario contracted for, when the harness says.

tests(timeout=300)[source]

Runs the suite the scenario declared, and reports the runner’s own summary.

Declared and never detected. The obvious detection is npm test, whose meaning is read from package.json - a file inside the perimeter the measured agent may edit, so broken code plus a test script of echo ok scores green. A detected command hands the choice of how a run is measured to the agent being measured, and no comparable tool has that adversary. Every documented migration elsewhere runs from detecting towards declaring and none the other way.

Three outcomes, not two. Green, red, and could-not-judge - the last covering an executable that is not there, a timeout, npm error Missing script, pytest’s exit 4 and 5, and node’s exit 7. Three distinct causes all exit 1, and calling any of them a failing suite would be scoring an agent for something it did not do.

No report format is parsed. One generic mechanism instead: find the summary the runner already wrote and hand it back, anchored on its marker and taken to the end. A table of four markers is acceptable where a table of four detections was not, and the asymmetry is the reason - a wrong marker degrades the reason while a wrong detection changes the measure.

Parameters:

timeout (int)

Return type:

Metric

probe(command, write=None, append=None, drop=None, timeout=30)[source]

Runs a behavioural probe against a copy of the measured tree.

A criterion that is a behaviour executes instead of being recognised: no pattern in the diff, no judge, no tokens, and a wrong answer is an assertion that breaks. This is the generic half of that; the probe’s own text and the cases it checks are the domain’s.

write creates or replaces files in the copy, append adds to the end of what is already there, drop removes what a glob selects. Appending is what replaces instrumentation by regular expression: a probe concatenated into the module runs inside the scope it measures, so it has nothing to enumerate. Measured against four trees on Node v26, that agrees with the regex everywhere the regex is right and works in four realistic cases where the regex silently finds nothing - a top-level class, a function*, a destructured declaration, and a collision moved into a new file. The regex therefore carried the very bias instrumentation exists to prevent.

Nothing generalises about visibility across languages, so the base holds no notion of it. What generalises is the placement rule above: the probe runs in the scope unit it measures - the module for JS and Rust, the package directory for Go, nothing needed for Python.

The copy is outside the clone, and that is not tidiness. The harness archives the diff after validation, and repo.diff runs git add -A –intent-to-add first, so an untracked file left in the clone enters the archived patch without condition and is replayed at every re-scoring as the agent’s own work.

Returns the probe’s parsed JSON. Raises CannotJudge when the probe could not be run or did not answer, and ProbeTimeout - a CannotJudge - when it ran too long. That default is the safe one; a domain that knows a slow probe means the agent’s fix loops forever can catch ProbeTimeout and score it as a failure, with a reason saying so. The choice belongs to whoever knows, not to the base.

Parameters:
Return type:

dict

sources_at_etalon(pattern, exclude=None)[source]

The text of the pinned files a pattern selects, joined.

Always from the tag, never from a working tree. A validator that fell back to the checkout’s working tree when the harness provided one was read as equivalent and is not: trysquare puts the source repository there, whose working tree is on main, so the reference drifts the moment main moves or someone fixes the issue in place - which is exactly what pinning by tag exists to prevent. Only the correct one is offered here.

The clone cannot serve either: it is made with –no-tags, so the tag does not exist in it. The source repository is the only thing that carries it.

Patterns are matched per path component, so game/*.js does not reach into game/sub/. The listing is not read again - files_at_etalon already holds it - so this costs one git show per selected file and nothing else.

The reading itself is repo.etalon_file, which the harness has had all along. Three validators written before this base existed reimplemented it with a raw subprocess.

Parameters:
  • pattern (str)

  • exclude (str | None)

Return type:

str

tool_calls()[source]

Every tool call of the run, in order, from the archived session.

issue1.py:386-450 reads context[“trace”], the raw stream, which is deliberately not archived (outputs.py:24-27: five hundred times the size, nothing the per-message record does not say). The calls were in the session all along, and in a better shape: one toolResult record carries the tool name, the id and isError together, so the toolCallId reconciliation that file needed had no cause but reading the wrong file.

One consequence worth stating: a metric of process is therefore replayable, because the session outlives the work directory. The only figure a session cannot give back is retries (measure.py:102-104).

Return type:

tuple[ToolCall, …]

property skills_expanded: tuple[str, ...]

The skills whose body the harness pasted into the prompt, in order.

The companion of tool_calls for the other half of one question. A scenario can load a skill two ways, and they leave opposite traces: loaded by name, the agent must read the SKILL.md itself, which is a call; referenced as /skill:<name> in the prompt, the body arrives already expanded and there is nothing left to read. A validator that knows only the call scores the second way as “the skill was never opened”, which is the opposite of what happened.

Named expanded and not loaded: every skill a variant declares is loaded, and this is the shorter list of those the agent did not have to ask for. Both are read from the archived session, so a metric built on either replays.

first_write(path)[source]

The index of the first call that wrote to path, or None.

The index rather than the call, because what a process metric asks is an ordering question: did the first write to the test file come before the first write to the source, with a failing suite between them. Two indices answer that; a list of calls makes every caller rediscover how.

Refuses rather than answering “no” when a call might have written and the base cannot tell - an unknown tool, or a subagent. See WRITES.

Parameters:

path (str)

Return type:

int | None

classmethod fake(**parts)[source]

An Assay for a test, answering only what the test declared.

Asking for anything else raises, and that is the point rather than a limitation. A fake that answered would put the absence of a measurement into the shape of a measurement - an empty set reads as “the agent touched nothing” - which is the confusion the error contract exists to prevent, moved into the tests. And a test that declares what it reads documents the dependency: a validator that grows a new read makes its old tests fail, loudly, which is information.

Return type:

Assay

trysquare.assay.validator(fn)[source]

Makes a plain function the whole of a validator.

fn stays callable, which is the point: a test scores a run with one call instead of a subprocess. test_issue1.py:40-54 had to ast.parse its own validator to list the metrics it produces, because calling it “would want a clone, a source repository carrying the tag and a trace” - that is what this replaces.

The if __name__ tail stays explicit rather than running at decoration time. A module that scored a run merely by being imported could not be imported by a test.

trysquare.outputs

The output tree, and the state that makes a matrix resumable.

One directory per experiment, and relaunching the same experiment overwrites it. The archive of previous versions is git. A timestamped directory per launch would accumulate variants of one experiment to choose between, which is optional stopping through the back door.

The directory name carries the experiment’s identity, which is also its guard: anything that changes what is measured changes the name, so a quick run at three repetitions writes to its own directory and cannot corrupt a published matrix at ten.

The layout, as a literal block so the underscores are not read as markup:

<output>/<scenario>_<etalon>_<provider>_<model>_n<N>/
  state.json        cells, runs, valid / empty / failed, attempt counters
  measures.json     one line per run
  synthesis.md      scores, costs, gaps and verdicts, written when complete
  runs/<cell>/<id>/           grouped, or runs/<id>/ when the tree is blind
    context.json  configuration.json  diff.patch
    session/*.jsonl          the agent's per-message record, one file per attempt
    validation/<mode>.json   validation/<mode>.stderr

runs/ takes one of two layouts, and which one is a property of the tree rather than of the measurement: a run id is an opaque hash, so a cell directory is the difference between reading a diff and looking a hash up first. Blind is the layout that costs something - a human filling a scoring form who knows which configuration they are grading grades it better - so a scenario declaring a form validator is blind by default and every other one is grouped. Output settles it once and is the only place a run’s path is built.

The layout is recorded in the state and deliberately not in the directory name: it changes where bytes land, not what is measured, so it is not part of the experiment’s identity.

The session files are the agent’s own trace, copied here so it outlives the work directory - which is disposable by design and which the OS may purge. The raw event stream is not copied: it is almost entirely streaming deltas and says nothing the per-message record does not, at five hundred times the size.

trysquare.outputs.LIVE = 'live.json'

it holds only runs in flight, and a finished matrix’s copy of it is empty by construction.

Type:

What the launch is doing while it is doing it. Not part of the archive

trysquare.outputs.write_json(path, payload)[source]

One serialization for everything this tree holds, and one write.

indent=2 and real UTF-8, with a final newline, declared once: a rewrite of the same data must be byte-identical to the original wherever it is written from, or replay –rescore could not promise to leave untouched what it did not change.

Written to a neighbour and renamed over the target, because state.json and measures.json are now rewritten after every run of a matrix that runs for hours. os.replace is atomic within a filesystem, so a kill during a write leaves the previous complete file rather than a truncated one - a ledger cut in half is worse than a ledger one run out of date, since nothing downstream can tell it is not the whole story.

Parameters:

path (Path)

Return type:

None

trysquare.outputs.write_text(path, text)[source]

The same neighbour-and-rename, for what is not JSON.

diff.patch is the one large plain write an archive makes, and a truncated patch is indistinguishable from a complete one: replay would reconstitute half of an agent’s work and say nothing about it.

Parameters:
Return type:

None

trysquare.outputs.slug(value)[source]

A value reduced to one path component.

Tags legitimately contain a slash (release/1.0), which would otherwise turn one directory name into two and put the result somewhere nobody named.

Return type:

str

trysquare.outputs.identity(scenario)[source]

The directory name without its repetition count.

What the name says about what is measured, as opposed to how many times. Two directories sharing it are the same experiment measured a different number of times, and run_id ignores the count, so the shorter one’s runs carry the very ids the longer one asks for: they are those runs, not analogues of them.

Return type:

str

trysquare.outputs.cell_dir(cell)[source]

One path component per cell.

A grid cell is named by joining its axis values with ` / , which `slug alone would turn into nothing—off. The separator becomes one underscore, as in the experiment’s own directory name.

Parameters:

cell (str)

Return type:

str

trysquare.outputs.run_location(runs_dir, run_id_, cell)[source]

Where one run’s directory is. The only rule, so both layouts have one author.

cell is None for a blind tree, and also for an id the scenario does not plan - a tree written under another scenario, whose runs have no cell here to be filed under and so stay at the root.

Parameters:
Return type:

Path

trysquare.outputs.is_run_dir(directory)[source]

Whether a directory is a run’s rather than a cell’s.

By what a run leaves behind, because a cell directory holds nothing but run directories. A run that produced nothing still archives its session, and one that failed before its diff still holds its validation, so no single marker covers them all - and a run with none of these is a directory with nothing in it to read.

Parameters:

directory (Path)

Return type:

bool

trysquare.outputs.sniff_layout(runs_dir)[source]

The layout a tree already has, read from its shape. None when it has none.

Read rather than assumed because a tree whose ledger was lost is still a tree somebody wants to render, and guessing wrong there does not fail loudly - it reports missing sessions for runs that are sitting on disk.

Parameters:

runs_dir (Path)

Return type:

str | None

trysquare.outputs.ledger_run_dirs(experiment, state)[source]

Where each run of a ledger sits, in the layout that ledger records.

For a reader that has the state in hand and wants the runs it names, rather than whatever directories happen to be there.

Parameters:
Return type:

dict[str, Path]

trysquare.outputs.archived_run_dirs(experiment)[source]

Every run directory under an experiment, in either layout.

The leaf name is the run id in both, so a caller can still identify a run by the directory it reads - which is how a re-scoring finds the row to rewrite.

Parameters:

experiment (Path)

Return type:

list[Path]

trysquare.outputs.experiment_name(scenario, repetitions=None)[source]

The directory name, which is the experiment’s identity.

Parameters:

repetitions (int | None)

Return type:

str

trysquare.outputs.run_id(scenario_name, cell, repetition)[source]

A short opaque id, stable for a given (scenario, cell, repetition).

Opaque so a form can be filled without knowing which cell is being scored, and stable so a resume can tell an absent run from one already done. The mapping back to the cell lives in state.json, deliberately not in the form.

Parameters:
  • scenario_name (str)

  • cell (str)

  • repetition (int)

Return type:

str

trysquare.outputs.cell_fingerprint(scenario, cell)[source]

What one cell declares, as one short digest.

The directory name is the experiment’s identity. This is a cell’s, one level down. Without it the ledger records that a run belongs to rule / high and nothing about what rule / high was, so editing a delta and resuming keeps the runs measured under the old declaration and completes the matrix with the new one: two configurations published under one name.

Taken over scenario.declared(cell), which is also what a run is built from, so what the ledger promises and what the agent received cannot drift apart.

Over the declaration, not over the bytes it points at: a context brick is fingerprinted by its path, and editing that file in place is not caught here. That is the choice etalon already makes - a tag up front, the commit it resolved to recorded per run in configuration.json - and it is what keeps resolve off the disk, which is what keeps a dry run free.

sort_keys, because reordering two keys of a TOML block must not cost a matrix.

Return type:

str

trysquare.outputs.per_cell(runs)[source]

How many runs a mapping of run ids holds for each cell, in first-seen order.

Parameters:

runs (dict)

Return type:

dict[str, int]

trysquare.outputs.measures_in(directory)[source]

The rows a matrix directory holds, empty when it holds none.

A free function because three callers read a matrix that is not their own: an experiment being carried from, and compare’s two sides.

Parameters:

directory (Path)

Return type:

list[Run]

trysquare.outputs.matrices(root, scenario)[source]

Every matrix of this experiment already under root, by repetition count.

Matched by reconstructing the name rather than by globbing it: a model name may legitimately carry [ or *, which a glob would read as syntax and silently fail to match. A directory without a ledger is not a measurement, so it is not offered.

Parameters:

root (Path)

Return type:

dict[int, Path]

class trysquare.outputs.Carried(directory, repetitions, entries, rows, trees, concurrency, timeout, etalon_commit, cells, chain, stranded, mismatch)[source]

What one carry would move out of a lower matrix of the same experiment.

Parameters:
directory: Path
repetitions: int
entries: dict
rows: list
trees: dict
concurrency: int | None
timeout: int | None
etalon_commit: str
cells: dict
chain: list
stranded: int
mismatch: tuple[str, ...]
trysquare.outputs.carryable(root, scenario, plan, repetitions)[source]

The runs a lower matrix of this experiment already holds, or None.

Reads, and never raises or writes: this is what a –dry-run announces, so it must cost nothing and it must be the same answer a real launch acts on.

One source, never two. A matrix assembled out of three measurement sessions is not something a reader could be expected to unpick, so the candidates are ranked and the best one carries alone: a source that agrees on what is measured beats one that does not, and then the largest count wins - which is both the cheapest carry and the one an operator would predict.

Parameters:
Return type:

Carried | None

class trysquare.outputs.Prior(kept, leftovers, source)[source]

What is already on disk that a launch has to decide about.

Parameters:
kept: int
leftovers: int
source: Carried | None
property offerable: bool

Whether there is a decision to take at all.

A source that disagrees on what is measured is not one: it cannot be carried whatever anybody answers, and the launch says why on its own.

trysquare.outputs.prior(output)[source]

What this launch is about to overwrite, carry or complete.

Reads and decides nothing, so what a launch would ask can be tested without a terminal and a question can be built from the same facts the plan is.

Parameters:

output (Output)

Return type:

Prior

class trysquare.outputs.Output(root, scenario, repetitions=None, grouped=None)[source]

The directory for one experiment, and everything written into it.

Parameters:
  • root (Path)

  • repetitions (int | None)

  • grouped (bool | None)

prepare()[source]
Return type:

Path

property layout: str

The layout as the state records it.

location(run_id_)[source]

A run’s directory, whether or not anything has been written to it.

Parameters:

run_id_ (str)

Return type:

Path

relative_run(run_id_)[source]

The same path, relative to the experiment directory, for a link or a form.

Parameters:

run_id_ (str)

Return type:

str

run_dir(run_id_)[source]
Parameters:

run_id_ (str)

Return type:

Path

plan()[source]

Every run this experiment expects, keyed by its opaque id.

Return type:

dict

cell_drift(previous)[source]

Where the scenario and an existing ledger disagree about the cells.

Two mappings of cell name to run count: the cells the scenario declares and the ledger does not know, and the cells the ledger holds and the scenario no longer declares. The directory name carries the scenario, the etalon, the agent and the repetition count - not the cells - so a scenario that grew a variant reuses the directory of the matrix already published, and nothing but this comparison can say so.

Read from the ledger as found. load_or_create_state fills in the ids it does not know, which is the behaviour described here and also what erases the evidence for it.

Parameters:

previous (dict)

Return type:

tuple[dict[str, int], dict[str, int]]

fingerprints()[source]

What every cell of this scenario declares, one digest each.

Return type:

dict[str, str]

measured_cells(state)[source]

Cells holding a run a resume can no longer relaunch.

Parameters:

state (dict)

Return type:

set[str]

changed_cells(state)[source]

Cells declaring something other than what their results were measured under.

Only cells that already produced one: while every run of a cell is still missing or empty, nothing was measured under the old declaration and the new one simply replaces it. The same line to_do draws, for the same reason.

A cell with no recorded fingerprint is not a changed cell. Ledgers written before fingerprints existed carry none, and inventing one now would record today’s declaration as the one those runs were measured under.

Parameters:

state (dict)

Return type:

list[str]

read_state()[source]
Return type:

dict

write_state(state)[source]
Parameters:

state (dict)

Return type:

None

write_live(payload)[source]
Parameters:

payload (dict)

Return type:

None

load_record(overrides=None)[source]

The load a launch runs under, as the ledger records it.

Concurrency and timeout are written down whatever their origin. They condition the retry count and therefore every cost column, so a matrix that does not record its own load cannot have its costs read.

The declaration and the flag stay two fields rather than one merged number: the synthesis header prints both, and merged they would no longer tell an experiment that asks for fifty from one whose operator asked for it that day.

Parameters:

overrides (dict | None)

Return type:

dict

initial_state(overrides=None)[source]

A fresh ledger, recording the load that produced it.

The repository is recorded by its logical name only. Where it actually came from - a directory, or a URL and the commit its tag pointed at - is written per run by runner.archive, because this method is called from resolve() during a –dry-run, and a field derived from the disk or the network would stop a dry run from being free.

cells records what each cell declares, so a later launch can tell a cell that was renamed from a cell that was rewritten. Pure computation over the parsed scenario, like everything else here, so a dry run stays free.

Parameters:

overrides (dict | None)

Return type:

dict

load_or_create_state(overrides=None)[source]

Reuses an existing ledger, or refuses when it is a different experiment.

A different repetition count is a different experiment and belongs in its own directory. Since the count is part of the directory name this should be unreachable, but a ledger that disagrees with its own directory is worth catching rather than trusting.

The load is the one thing a reused ledger restates rather than keeps: it is not a result, it is what this launch is about to run at. Inherited, it would name the load of whichever launch created the directory, while the state, the synthesis header and compare’s line of declared differences all read it as the load the runs beside it were measured under.

Parameters:

overrides (dict | None)

Return type:

dict

replayed(state, cells)[source]

The ledger these cells’ results are discarded from. Writes nothing.

What –overwrite CELL means, one cell at a time instead of the whole matrix: a run of a named cell returns to MISSING and is measured again, and every other run of the ledger is left exactly as it was found.

attempts returns to zero and detail goes, because both describe the measurement being discarded. So does the per-run carried flag: a run this launch measures itself is this matrix’s own run, whatever matrix first paid for it.

Their fingerprints are re-recorded, and that is the half load_or_create_state cannot do: it freezes the declaration of every cell that produced a result, so a cell re-measured under a new declaration would keep the digest of the old one and the next –resume would refuse the very runs this launch just measured. Only the named cells - a digest written for a cell nobody re-measures would claim today’s declaration is the one its results were measured under.

Parameters:
Return type:

dict

seed(state, carried)[source]

The ledger this matrix would have with a lower one’s runs in it. Writes nothing.

Shared by resolve, which must leave the disk untouched, and by absorb, which is the write. A carried run keeps its own state, usage and attempts: nothing is re-measured, so nothing may be restated - and to_do then has nothing to relaunch for it, which is the whole mechanism.

The record is a list, appended to the source’s own, so an experiment extended twice can still be read back to the launch that first measured each run.

Parameters:
Return type:

dict

absorb(carried, overrides=None)[source]

Copies a lower matrix’s runs in, and writes them down as this one’s.

The trees are copied, not linked. replay –rescore rewrites validation/<mode>.json in place, and a link would make a re-scoring of this matrix silently rewrite the matrix it was carried from - which is the one that has to stay intact for the two to be compared at all.

Over load_or_create_state and never initial_state: a run this matrix measured itself must survive the carry. That is the same rule as everywhere else here - a result already paid for is out of reach.

Parameters:
Return type:

None

to_do(state, only=())[source]

The runs a launch should perform.

A valid run is never relaunched, whatever its result. That is the whole protection: a resume has no power over anything that produced a result.

A run of a cell the scenario no longer declares is not one to perform: there is nothing to launch it as. Launching it anyway spent a clone to reach unknown cell, recorded the run as empty, and left it resumable - so it came back on the next pass, and every pass after that.

Parameters:
Return type:

list[tuple[str, dict]]

record(state, run_id_, run)[source]
Parameters:
Return type:

None

summarise(state)[source]

What the ledger holds, and whether this scenario’s matrix is finished.

Counted over the cells the scenario declares. A cell it no longer declares - a variant renamed between two launches - keeps its runs and keeps them rendered, but its unfinished ones say nothing about the matrix being planned, and counting them held complete at false with no launch able to lift it.

Parameters:

state (dict)

Return type:

dict

write_measures(runs)[source]
Parameters:

runs (list[Run])

Return type:

Path

read_measures()[source]
Return type:

list[Run]

archive_sessions(run_id_, session_dir, exclude=frozenset({}))[source]

Makes a run’s session archive equal to what this launch produced, as jsonl.

Two filters, and both exist to keep one launch’s traces from being read as another’s.

exclude names session files the caller does not want, by file name. The work directory keeps a run’s session directory from one launch to the next - the run id is stable, so the path is - so copying whatever is there would import a previous measurement’s traces.

And the archive is replaced, not added to. Relaunching an experiment overwrites it, so a session left by the previous launch would be attributed to this one: the file count would stop matching attempts, and a page rendered from the old trace would sit there looking current.

An absent session_dir is not an error: the agent may have failed to start at all, and that is a run to record, not an exception to raise.

Parameters:
Return type:

list[Path]

sessions(run_id_)[source]

A run’s archived sessions, in order. Empty when none were archived.

Parameters:

run_id_ (str)

Return type:

list[Path]

write_synthesis(text, suffix='')[source]
Parameters:
Return type:

Path

write_configuration(run_id_, configuration)[source]
Parameters:
  • run_id_ (str)

  • configuration (dict)

Return type:

Path

write_validation(run_id_, mode, payload, stderr='')[source]
Parameters:
Return type:

None

trysquare.outputs.incomplete_note(counts)[source]

The line that keeps an incomplete matrix from reading as a result.

Parameters:

counts (dict)

Return type:

str

trysquare.outputs.carried_note(state)[source]

The paragraph that keeps an extended matrix from reading as one launch.

Runs are interleaved so that the durations of one matrix are comparable between its cells, under one provider load (see runner). A carried run was measured under another, so an extended matrix holds two sessions in its cost columns.

Which is a legitimate thing to publish and not a legitimate thing to leave unsaid, so it is written where a reader who never saw the terminal will meet it - beside table.retry_warning, which exists for exactly this class of reservation.

Parameters:

state (dict)

Return type:

str

trysquare.outputs.unmeasured_note(state, runs)[source]

The line for runs the ledger counts and measures.json has no row for.

A matrix written before measures were saved run by run can hold this: the runs are on disk, the ledger calls them measured, and no table can see them. Publishing then reports a smaller n than was paid for, under a header naming the full repetition count, which is the whole family of defect this tool exists to refuse.

Nothing here can repair it - a row is derived from a measurement, and re-scoring needs a row to start from - so the honest move is to name the runs and say what getting them back costs.

Parameters:
Return type:

str

trysquare.runner

Orchestrating a matrix: what runs, in what order, and what is written down.

Runs are interleaved across cells rather than grouped by cell. That is not a scheduling detail: interleaved runs see the same provider load, which is the only reason durations are comparable between cells of one matrix. They are never comparable between matrices, and nothing here pretends otherwise.

An exception in one run must not cost the matrix. A missing cell beats a lost table, and the runs already paid for are the ones being protected.

class trysquare.runner.Plan(scenario, config, output, repo_path, repo_source, todo, overrides, blindness, notes, carried=None, replay=())[source]

Everything settled before a single token is spent.

Parameters:
scenario: Scenario
config: Config
output: Output
repo_path: Path
repo_source: str
todo: list[tuple[str, dict]]
overrides: dict
blindness: dict
notes: list[str]
carried: Carried | None = None
replay: tuple[str, ...] = ()
property runs: int
load(name)[source]

What this launch actually runs name at: concurrency, timeout, attempts.

The one place the precedence is spelled out, so the forecast an operator reads before spending, the pool that spends it and the ceiling each run is stopped at cannot answer differently.

The flag wins because it is explicit, announced in the header and recorded in the ledger; then the scenario, which is the experiment; then the machine, which may supply a fallback and never a measurement.

Distinct from what the ledger records, which keeps the declaration and the flag as two fields rather than the one number they resolve to.

Parameters:

name (str)

trysquare.runner.resolve(scenario, config, output_root, overrides=None, only=(), resume=False, grouped=None, extend=False, replay=())[source]

Turns a scenario into a concrete plan, refusing what cannot be measured.

Parameters:
Return type:

Plan

trysquare.runner.refuse_unknown_cells(flag, given, names)[source]

Refuses a cell name no cell of the scenario answers to, whichever flag gave it.

The refusal lists every cell, because a grid names its cells by joining axis values and the exact spelling is easier to copy than to guess.

Parameters:
Return type:

None

trysquare.runner.refuse_changed(output, changed, replay)[source]

Why a launch that keeps results cannot keep these ones.

Two wordings for one defect, because the way out differs: a resume keeps every cell, so any of them may be the one rewritten; a replay already discards the cells it names, so a cell reported here is one it was told to keep. Both end on the same offer - measure those cells again, and keep the rest of the matrix - which is what –overwrite CELL is for.

Parameters:
Return type:

str

trysquare.runner.available_note(carried)[source]

What a launch says about a lower matrix it is not carrying.

The only place –extend is discoverable at the moment it would be useful. An operator who has just asked for twenty repetitions is about to pay for ten they already own, and nothing else in the output would tell them.

Parameters:

carried (Carried)

Return type:

str

trysquare.runner.refuse_carry(carried, root, scenario, repetitions)[source]

Why –extend has nothing to work with, naming what is actually on disk.

Refused rather than quietly measured from scratch, for the reason –only with a typo is refused: a flag that silently does nothing looks like a flag that worked.

Parameters:
Return type:

str

trysquare.runner.drift_notes(output, previous)[source]

What an existing ledger and the scenario disagree about, before the first token.

Adding a cell to a published matrix already works: the ledger gains the new ids as missing, and a resume relaunches only what produced nothing. Nothing said so, and a saving nobody announced is a saving nobody takes - the safe-looking move was to remeasure a matrix that was already paid for.

Neither note depends on –resume: both state what a resume would do, so they are as true when the flag is absent as when it is there.

Parameters:
Return type:

list[str]

trysquare.runner.settle_repo(scenario, config)[source]

Where the repository will be read from, and what the config declared.

A repository entry may be a directory or a URL. Either way this only computes where it will be read from; nothing is cloned or reached for until execute.

Parameters:
Return type:

tuple[Path, str]

trysquare.runner.refuse_unmeasurable(scenario)[source]

The refusals that spend nothing, shared by resolve and validate.

Every brick must be where the scenario says, checked before anything is spent. A missing brick used to surface as a validator failure after a matrix had been paid for, or worse, as a prompt path silently sent to the agent as literal text.

Parameters:

scenario (Scenario)

Return type:

None

trysquare.runner.interleave(todo)[source]

Orders runs by repetition first, so cells are measured side by side.

Grouping by cell would measure the first cell under an idle provider and the last under a loaded one, and the duration column would then compare our own scheduling rather than the configurations.

Parameters:

todo (list[tuple[str, dict]])

Return type:

list[tuple[str, dict]]

trysquare.runner.brick_paths(scenario, config, cell, base)[source]

The bricks one cell loads, resolved to real paths.

Parameters:
Return type:

dict

trysquare.runner.looks_like_path(value)[source]
Parameters:

value (str)

Return type:

bool

trysquare.runner.source_dir(config, name, url, etalon)[source]

Where a remote repository is pinned. A pure path computation, no disk.

Keyed by tag, like harness/{name}-{tag}, and that is what deletes the entire staleness question: a directory already there is by construction already at the tag being asked for. Nothing to refetch, nothing to verify, and a tag moved upstream cannot leak into a matrix in flight. The cost is re-cloning when a matrix changes etalon, which is the right trade against a cache-invalidation state machine.

Keyed by a hash of the URL as well, so editing the URL in the config lands somewhere else instead of silently reusing the previous repository’s clone - the same class of defect as reusing a half-written harness.

Parameters:
Return type:

Path

trysquare.runner.prepare_source(config, name, etalon)[source]

The local repository runs clone from, pinning a remote exactly once.

Called from execute before anything is written, so an unreachable URL costs a refusal and an untouched disk instead of a ledger full of empty measures; and again from one_run, so no caller can reach a clone without having pinned first. On the hit path that second call is one lock and one is_file() against a run that lasts minutes.

Serialised for the reason prepare_harness documents: cells run concurrently and every one of them arrives here at the same moment. The readiness marker records the URL as well as the tag, so a directory named after a hash can say what it is, and a marker left by a different URL forces a fresh clone rather than being trusted.

Parameters:
Return type:

Path

trysquare.runner.prepare_harness(config, name, tag)[source]

Clones a pinned harness repository once, and installs it, exactly once.

Serialised, because cells run concurrently and every cell that loads this brick reaches here at the same moment. Checking exists() and then cloning is not atomic: the first real run of a scenario with four concurrent cells had two of them fail on destination path already exists, and a third succeeded only by winning the race - which meant it may have loaded an extension whose dependencies were never installed.

A readiness marker rather than mere existence, so a clone left half-written by an interrupted attempt is redone instead of silently reused. Reusing a partial harness is the kind of failure that produces a plausible measurement.

A [harness] entry may be a URL, exactly like a [repos] entry: the two sections have the same shape, and one accepting an address the other refuses would be an asymmetry nobody could guess. Nothing clones from this directory - it is loaded as an extension - so it keeps –no-tags.

Parameters:
Return type:

Path

trysquare.runner.install_dependencies(clone_dir, timeout=300)[source]

Installs a harness repository’s runtime dependencies, once per pinned clone.

A freshly cloned extension is not loadable as it stands: pi-subagent declares @earendil-works/pi-coding-agent as a runtime dependency, and the agent resolves imports from node_modules next to the extension or above it. Without this the extension fails to load, and a cell that was supposed to measure a toolkit measures its absence.

A missing package.json is not an error: a brick may be a single file.

Parameters:
Return type:

None

trysquare.runner.read_brick(base, value)[source]

A scenario value that may be inline text or a path to a file.

A value that looks like a path and does not exist raises. It used to fall back to being treated as inline text, and that silent fallback is exactly the defect this project keeps paying for: a mistyped prompt path became the literal string “tickets/vague.md” sent to the agent as its task, and the runs looked entirely normal while measuring nothing.

Parameters:
Return type:

str | None

trysquare.runner.referenced_paths(scenario, base)[source]

Every file the scenario points at, with where it was declared.

Collected so a missing brick is refused before the first token, not discovered by a validator failure after a matrix has been paid for.

Parameters:
Return type:

list[tuple[str, Path]]

trysquare.runner.preflight(scenario, base)[source]

Refuses a scenario whose bricks are not where it says they are.

Parameters:
Return type:

list[str]

trysquare.runner.one_run(plan, run_id, meta, board=None)[source]

Measures one cell once, and writes down everything about it.

Every failure path here ends in a Run with a state, never in an exception: one frozen run must not take the matrix with it. A stop is the exception to that, and interrupt.Stopped is shaped so it cannot be caught here - see below.

Parameters:
Return type:

Run

trysquare.runner.judge(plan, validator, directory, base, prompt, response, clone, attempts, work)[source]

Assembles the judge’s dossier and runs it.

The pieces are whatever the scenario declared, and nothing else - certainly not the cell. response is the agent’s final prose rather than its transcript: a judge asked whether a note is usable must score the note, not the work behind it.

work is where the judge’s own stream goes. It is passed rather than derived from directory, which lies inside the published archive: a 16 MB stream per run has no business in a tree meant to be committed, and –extend copies that tree whole.

Parameters:
trysquare.runner.stream_ceiling(plan)[source]

How many bytes one agent run may write, from a config expressed in megabytes.

Megabytes in the file a human writes, bytes at the one place that compares a size, so the unit lives where it is read rather than in every signature it passes.

Parameters:

plan (Plan)

Return type:

int

trysquare.runner.recorded_model(sessions)[source]

The model the archived sessions say answered, or None when they do not say.

None rather than the declared pattern: an archive that cannot name the model must say so, and filling the gap with the intention is how a fallback would hide.

Parameters:

sessions (list[Path])

Return type:

str | None

trysquare.runner.archive(plan, run_id, clone, prepared, cell, thinking)[source]

Keeps the sources a re-score needs, and nothing more.

The raw stream is almost entirely streaming deltas, which teach nothing the per-message record does not: 15.9 MB of stream against 30 KB of session. What is archived is the tag, the diff and the configuration, which is exactly what replay needs to reconstitute a tree.

The repository and the commit its tag resolved to are recorded here. Without them a published archive cannot say what it measured - and with a URL the address is the only thing that identifies it. The commit also closes a hole that predates remotes: a local repository whose tag was moved between two matrices left no trace at all.

model_id is that same distinction one level up. model is a pattern the agent resolves against what the provider offers - gemma-4 runs as gemma-4-31b - so the declared value is an intention and only the session says what answered. Keeping the pattern alone left the archive unable to name the model it measured, and made a fallback to the machine’s defaultModel indistinguishable from a resolution.

Parameters:
Return type:

None

trysquare.runner.unconsumed(futures, already)[source]

The runs that had finished while the loop was not looking.

as_completed hands runs over one at a time, so an interrupt in the loop body abandons every run that finished behind the one being written down. They were paid for at the same price as the one that got recorded.

The order of the tests is not free. exception() raises on a cancelled future, so cancellation is checked first; and a future carrying a Stopped is a run that was cut short, which is the one thing that must never be written down.

Parameters:
Return type:

list[Run]

trysquare.runner.execute(plan, on_run=None)[source]

Runs the plan, writing state and measures as it goes so an interruption is resumable.

The two are written together, run by run - see keep, which is also where the order of the two writes is argued. Writing the ledger alone was enough to resume and not enough to keep what had been paid for: a Ctrl-C left runs marked valid in state.json with no row in measures.json, and those runs were then out of reach - –resume relaunches only what produced nothing, and replay has no row to re-score. The matrix went on to publish as complete over fewer runs than were measured, and nothing in the output said so.

An interrupt keeps everything that was finished and records nothing that was not: what had completed unseen is harvested on the way out, what was still running is left missing for the next –resume.

The repository is pinned first, before a single directory is created. After output.prepare() an unreachable URL would leave behind an experiment directory holding a ledger of runs that never had a repository to run against; before it, the refusal reaches the operator and the disk is as untouched as after a dry run.

Parameters:

plan (Plan)

Return type:

list[Run]

trysquare.progress

A bar for the loops that take hours, and an honest estimate of what is left.

A matrix is dozens of agent runs of several minutes each. Until now the only thing an operator saw was one line per finished run, which says what was measured and nothing about what is left: no count, no elapsed time, no arrival estimate. On a matrix that runs overnight, “still working” and “hung” looked the same from the terminal.

Three rules hold this together.

The record scrolls, the bar is pinned. Every line a command printed before still prints, unchanged, above the bar. Bar.line exists so a caller never has to know whether a live region currently owns the terminal.

The estimate is the average rate since launch, not a recent one. See eta_seconds.

Off means off. When output is not a terminal - a pipe, a redirect, a test capturing stdout - there is no bar and no escape sequence, and Bar.line is a plain print. That is what keeps a piped run byte-identical to what it printed before this module existed.

trysquare.progress.OFF = 'TRYSQUARE_NO_PROGRESS'

Turns the bar off without a command line. For wrappers and CI, which cannot always reach the arguments of the command they run.

trysquare.progress.eta_seconds(completed, total, elapsed)[source]

Seconds left, from the throughput since launch. None until it means one.

Rich’s own TimeRemainingColumn divides by Task.speed, which averages completions inside a thirty-second window. A matrix of ten-minute runs completes nothing at all inside most thirty-second windows, so that column would read -:–:– for the whole matrix except the instants a batch lands.

Throughput since launch is what an operator computes by hand, and on a saturated pool it is exactly right rather than approximately: five runs of T done out of thirty-two gives 27 x T/5, which is what the remaining twenty-seven will take five at a time. It counts the ramp before the first completion instead of discarding it, and it steps down in batches rather than swinging by minutes every few seconds.

It over-estimates only at the tail, where fewer runs are left than there are workers. An estimate that arrives early is the one to prefer.

Parameters:
Return type:

float | None

trysquare.progress.clock(seconds)[source]

A duration at the scale it is read at.

A matrix runs for hours, so 1h 12m is the useful reading and 72m is not. Both the elapsed and the remaining column use this, so the two read alike.

Parameters:

seconds (float)

Return type:

str

trysquare.progress.wanted(stream=None, no_progress=False)[source]

Whether a bar should be drawn at all.

NO_COLOR is deliberately not read: it asks for no colour, not for no motion, and rich already honours it for styling.

Parameters:

no_progress (bool)

Return type:

bool

class trysquare.progress.Bar(progress=None, task_id=None)[source]

A counter being advanced, and a print that survives it.

Constructed with progress=None when there is no terminal to draw in, so a caller has one code path either way.

Parameters:

progress (Progress | None)

property enabled: bool
tick(step=1)[source]

Advances the bar. Past the total is allowed: a resumed matrix can fire its callback more often than it planned to, and that must not raise.

Parameters:

step (int)

Return type:

None

line(text='')[source]

Writes one line above the bar, or straight out when there is no bar.

Parameters:

text (str)

Return type:

None

warn(text)[source]

The same, for what belongs on stderr.

It only reaches stderr when the bar is off. A live region owns the terminal it draws in, and a second stream writing into that region tears it; the bar is only ever on when stdout is a terminal, and then both streams are that same terminal anyway. Redirect or pipe either one and the bar is off, which is the case where the distinction can still be observed.

Parameters:

text (str)

Return type:

None

trysquare.progress.bar(total, label, enabled=True)[source]

A bar pinned to the bottom for the length of the with, or nothing at all.

Nothing at all is a Bar too. And it is a context manager because an interrupt during a matrix is normal: Progress restores the cursor on the way out, before the entry point catches the interrupt and prints what survived.

One exit skips every finally there is, and it is the one the terminal cannot afford: interrupt leaves with os._exit when a child will not answer a signal. on_hard_exit is how the live region is given back on that path too. Registered inside the with, because before it there is no region to give back.

Parameters:
Return type:

Iterator[Bar]

trysquare.interrupt

Stopping, when the operator asks.

A matrix runs for hours, so interrupting one is normal rather than exceptional. Two things made it neither quick nor clean, and both come from the same absence: nothing owned the processes the harness starts.

Nothing stopped. ThreadPoolExecutor shuts down with wait=True, and CPython puts the shutdown sentinel behind every pending work item, so the whole queue still ran. A matrix of thirty-two runs stopped at the fifth spent another hour, with the bar frozen and nothing printed. Cancellation has to be cooperative - concurrent.futures joins its workers at interpreter shutdown, so not even sys.exit escapes - and cooperative cancellation usually means a flag tested in a dozen places, one of which is always forgotten.

One door instead of a dozen checks. Every subprocess in this package goes through run, and run refuses to open once the operator has asked to stop. A queued run therefore dies at its first git clone without anyone having written a check for it, and a run already in flight dies when its child does. The one flag test outside this module is at the top of runner.one_run, and it exists only so a worker released from a lock does not get as far as creating a directory.

Children are session leaders. start_new_session=True is set here, not at the call sites, because it is inseparable from the killing: os.killpg then reaps the whole descendant tree, which is what it takes to stop a validator’s pytest or an agent’s own tool subprocesses. It also removes an accident this used to rely on. Children shared the terminal’s process group, so a Ctrl-C reached them for free - and only a Ctrl-C did. A kill, a CI cancellation, a docker stop reached the harness alone and orphaned every agent to burn tokens for a full timeout with nobody watching. Now every route is the same route.

Asking, then taking. The first signal asks: the flag goes up, every child gets SIGTERM, and Stopped unwinds the main thread. A watchdog then guarantees an end even if a child ignores signals - SIGKILL at GRACE, and the process leaves at DEADLINE. A second signal takes that path at once. This is the whole reason the escalation is armed by the handler rather than by stop: a stop called while the main thread is already unwinding needs no deadline.

trysquare.interrupt.GRACE = 2.0

How long a child gets to answer SIGTERM before it is killed outright.

trysquare.interrupt.DEADLINE = 5.0

How long the whole shutdown gets before the process leaves without it.

trysquare.interrupt.POLL = 0.25

How often a child under a ceiling is weighed. What it writes between two weighings is what it may overrun by, so this is a fraction of a second rather than one.

exception trysquare.interrupt.TooMuchOutput(what, limit)[source]

A child stopped for writing more than the caller allowed.

An exception rather than an exit status, for the reason run gives about the status of a child we killed: it is not evidence about anything. An Exception rather than a Stopped, because one runaway run must be written down in the ledger, not unwind the matrix around it.

Parameters:
Return type:

None

exception trysquare.interrupt.Stopped(signum=None, message='')[source]

A subprocess that the operator’s interrupt cancelled, raised instead of run.

A KeyboardInterrupt rather than an Exception, and that is load-bearing rather than decorative. runner.one_run wraps a whole measurement in except Exception, so that one frozen run cannot cost the matrix, and every failure path there ends in a Run carrying a state. A cancellation caught by it would be written down: a run interrupted after the agent produced tokens keeps valid, and valid is not in outputs.RESUMABLE, so no later –resume could reach it again - a run with no metrics and no diff, recorded as measured forever. Passing through except Exception untouched is what leaves a cancelled run unrecorded, and therefore still missing for the resume that follows.

Parameters:
  • signum (int | None)

  • message (str)

Return type:

None

trysquare.interrupt.stopping()[source]

Whether the operator has asked for this to end.

Return type:

bool

trysquare.interrupt.signalled()[source]

The signal that asked, when one did. None for a stop asked for in Python.

Return type:

int | None

trysquare.interrupt.reset()[source]

Back to a state where work may start, and any armed watchdog disarmed.

For tests. A process that has been asked to stop does not resume.

Return type:

None

trysquare.interrupt.on_hard_exit(callback)[source]

Registers something to do before the process leaves without unwinding.

For what owns the terminal rather than what owns a file: os._exit skips every finally, so a live region abandoned mid-draw leaves the terminal with no cursor. Files need nothing here - outputs.write_json writes a neighbour and renames it, so a hard exit during a write leaves the previous complete file.

Return type:

None

trysquare.interrupt.stop(signum=None)[source]

Refuses every child from here on, and takes down the ones already running.

Idempotent, which is what lets the signal handler, the pool and a caller unwinding all say it without coordinating.

Parameters:

signum (int | None)

Return type:

None

trysquare.interrupt.run(argv, *, cwd=None, env=None, input=None, timeout=None, capture_output=False, text=False, stdin=None, stdout=None, ceiling=None, written=None)[source]

subprocess.run, for a process the harness can still stop.

A reimplementation rather than a wrapper, because subprocess.run gives no way to reach the Popen it creates, and reaching it is the entire point.

Two refusals, and the second is the one that is easy to leave out. Before the spawn, so a queued run dies at its first child instead of measuring anything. And after the wait, because a child we killed comes back as returncode -15 and every caller would read a verdict into it: an unproductive agent worth retrying, a validator exited -15, a could not install dependencies. An exit status we caused is not evidence about anything.

stdout is a file the caller opened, and it overrides the pipe capture_output would have given: communicate accumulates a pipe in the parent, and the agent’s stream has no size anyone controls. Handing the child a file is the only way a caller can bound its own memory. capture_output then still means stderr, which is small and which the caller needs as a value. CompletedProcess.stdout is None: the file is the output, on a timeout too, where the exception carries nothing because there was no pipe to carry it from.

ceiling bounds what the child produces the way timeout bounds the wait, and for the same kind of child: one that will not end on its own. written says how much that is, and answering it is the caller’s job - a file holds every byte, a pipe may have a sieve at its far end keeping a thousandth of what arrives, and which of the two is being bounded is not a decision this module can make.

Two mechanisms, and neither replaces the other. A thread weighs the output while the child runs and kills its group past the ceiling, which is what stops a runaway from taking the machine and the runs beside it. The verdict is then asked again after the wait, because a child fast enough to overrun between two weighings would otherwise be accepted for having been quick about it.

The timeout branch does what subprocess.run does and nothing more. For a piped caller the exception was built by Popen._check_timeout before _communicate reached its text-mode translation, so its partial output is bytes even under text=True. Collecting the streams again here would quietly break that.

Parameters:
Return type:

CompletedProcess

trysquare.interrupt.handled(*signums)[source]

Answers the operator’s signals for the length of the with, then gives them back.

Off the main thread this installs nothing and says so by doing nothing: signal.signal refuses there, and raising would make cli.main uncallable from a test that runs it in a thread.

Parameters:

signums (int)

Return type:

Iterator[None]

trysquare.interrupt.hard_exit(signum=None)[source]

Leaves now, having given the terminal back first.

os._exit skips every finally, every atexit hook, and the join concurrent.futures performs on its workers - which is exactly why it is here, since a worker wedged on a child that ignores signals is the only way to reach this. What it must not skip is whatever owns the terminal.

Parameters:

signum (int | None)

Return type:

None

trysquare.cli

The command line.

Eight subcommands, and –output roots every one of them that writes.

Overrides are always announced at launch. That is not politeness: the previous tool had a protocol declared in a document and defaults in the code that contradicted it, and the code wins at the moment somebody types the command, so a published matrix was measured at the wrong load. A plan that cannot be executed is not a plan.

Overrides are then stamped according to their effect. Anything that changes what is measured goes into the directory name, so a quick run at three repetitions writes elsewhere and cannot corrupt a published matrix at ten. Anything that changes the load goes into state.json and the synthesis header, because it conditions the retry count and therefore every cost column.