# Outputs Everything is rooted at `--output`. One directory per experiment. ```text /____n/ state.json the per-run ledger measures.json one line per run synthesis.md the score table, the cost table, the gap table and the verdicts synthesis.html the same synthesis as one self-contained page runs/// configuration.json what this run actually ran diff.patch what the agent changed session/.jsonl the agent's own trace, one file per attempt session/.html the same trace as a page, on `render --html` validation/.json each validator's output validation/.stderr kept when a validator fails validation//context.json what that validator was handed ``` `context.json` lives under its validator, not at the run's root: every validator gets its own, and a judge's is blinded where a script's is not. One file at the root would have to be two files under one name. (runs-layout)= ## Two layouts for `runs/` A run id is an opaque hash, so a cell directory is the difference between reading a diff and looking a hash up first. ```text runs/ runs/ rule_off/ 658df337/ 658df337/ 962d7594/ 962d7594/ 1af14a46/ nothing_off/ 1af14a46/ ``` Grouped on the left, blind on the right. **Grouped is the default**, and blind is the default for a scenario that declares a `form` validator - there, the id is what keeps a human from knowing which configuration they are grading, and somebody who knows they are scoring the best-equipped cell scores it better. `run --group-by-cell` / `--no-group-by-cell` settles it either way. Asking for a layout that contradicts a tree already on disk is **refused**: the two do not merge, so every run would end up archived twice, once under each, with nothing to say which half was this launch. The layout is recorded in `state.json` and deliberately **not** in the directory name: it changes where bytes land, not what is measured. Every later command - `render`, `replay`, `form`, `parity --smoke` - reads it from the tree, so none of them needs the flag repeated. The leaf is the run id in both layouts. That is what the ledger, the measures, the sessions and a scoring form all name a run by, and one identity for one run is what lets a re-scoring find the row it belongs to. ## The directory name is the guard The name carries the experiment's identity: scenario, etalon, provider, model and the repetition count. That is not decoration. Anything that changes what is measured changes the name, so `--repetitions 3` writes to `..._n3/` and **cannot** overwrite a published matrix at ten. Relaunching the same experiment **overwrites** it. The archive of previous versions is git. A timestamped directory per launch would accumulate variants of one experiment to choose between, which is optional stopping through the back door. ## `state.json` The ledger, and what makes a matrix resumable. ```json { "scenario": "rule-vs-ticket", "etalon": "etalon-v1", "provider": "ilaas", "model": "gemma-4-31b", "thinking": "off", "repetitions": 2, "concurrency": 5, "timeout": 900, "layout": "by-cell", "overrides": { "repetitions": 2 }, "complete": true, "cells": { "nothing / off": "b6f0c2a1d4e37f58", "rule / off": "0a91c7d5e2b48366" }, "runs": { "658df337": { "cell": "nothing / off", "repetition": 0, "state": "valid", "attempts": 1 }, "962d7594": { "cell": "nothing / off", "repetition": 1, "state": "empty", "attempts": 2 } } } ``` Four states: `missing`, `empty`, `validator_failed`, `valid`. Only `missing` and `empty` are **resumable** - the two that produced no result at all. A `valid` run is never relaunched whatever its result, and `validator_failed` is re-scored rather than re-measured. `attempts` accumulates across resumes, so an abusive one leaves a trace. The ledger holds exactly the cells the scenario declared when it was written, and a later edit can make the two disagree. `run` compares them before spending, and says which way: - a cell the scenario declares and the ledger does not know is **added** to it, its runs `missing`, so `--resume` measures that cell and leaves every measured run alone. This is how a variant is added to a matrix already published - see {doc}`../guide/writing-a-scenario`. The plan announces it as `ADDED:`. - a cell the ledger holds and the scenario no longer declares is **kept**. Its runs stay in every table, because a cell must never vanish from a synthesis silently. None of them is launched or counted towards `complete`: the scenario has no such cell to run, so an unfinished one would hold the matrix incomplete with nothing able to lift it. The plan announces it as `STALE:`, and the answer is to delete the directory rather than resume onto it. `cells` is what makes the third case catchable: a cell **rewritten** under its own name. The directory name guards the experiment; this guards each cell inside it. The digest covers what the cell declares - its own `prompt`, `context`, `system` and `thinking`, the `[task].prompt` and `[agent].thinking` it falls back to, and the `[harness]` entries it names. A cell that has produced nothing takes the new digest; one that has produced a result is frozen, and a `--resume` that no longer matches it is refused. It is a digest of the **declaration**, not of the bytes it points at: a context brick is covered by its path, and editing that file in place is not caught. Same choice as `etalon` - a tag up front, the commit it resolved to recorded per run in `configuration.json`. A ledger written before digests existed carries none, and an absent digest is not a changed one. The load is recorded whatever its origin, because retries depend on it and every cost column depends on retries. ### `carried`, when a matrix was extended Absent from every matrix measured in one launch. `run --extend` adds it, and a run that travelled is marked `"carried": true` beside its state: ```json "carried": [ { "from": "rule-vs-ticket_etalon-v1_ilaas_gemma-4-31b_n10", "repetitions": 10, "runs": 60, "concurrency": 5, "timeout": 900, "etalon_commit": "a1b2c3d" } ] ``` A **list**, so an experiment extended twice can still be read back to the launch that first measured each run. `etalon_commit` comes from the carried runs' own `configuration.json`: the ledger records the repository by its logical name only, so a tag moved upstream between two launches would otherwise leave no trace anywhere. Nothing here is carried silently - the launch says so, and so does the synthesis, header line and `:warning:` paragraph both. :::{note} `state.json` also holds the **id-to-cell mapping**, which is deliberately *not* in the manual scoring form: the form is blind, and so is the tree it points into whenever a `form` validator declares that a human scores. See [Two layouts](#runs-layout). ::: ## `measures.json` One entry per run, and the raw material every table is rebuilt from. ```json [ { "id": "658df337", "cell": "nothing / off", "repetition": 0, "usage": { "input": 14036, "output": 2286, "turns": 6, "retries": 2, "cost": 0.0 }, "duration": 64, "metrics": { "overflow": true, "delivered": true, "tests": true, "issues": ["#1"] }, "reasons": { "overflow": "addressed without being asked: #1" }, "state": "valid", "attempts": 1 } ] ``` Persisting **per-run** values rather than aggregates is what makes `render` possible. Keeping only medians would make a matrix permanently unusable for a verdict, and a matrix costs hours. ## `synthesis.md` Written **only when the matrix is complete**. An incomplete matrix says so instead: ```text This matrix is incomplete: 3 never launched, 1 produced nothing. No synthesis is published; `--resume` completes it. ``` Three tables, in this order. The **score table** says what each cell did: cells in rows, in the order the scenario declares them, and one column per declared boolean metric. ```text | cell | overflow | delivered | in_scope | tests | | -------------------- | -------- | --------- | -------- | ----- | | nothing / off | 10/10 | 10/10 | 9/10 | 9/10 | | rule / high | 2/10 | 10/10 | 8/9 | 9/10 | | careful ticket / off | 0/10 | 10/10 | 9/10 | 9/10 | ``` `x/n` counts the runs where the test was true out of the runs that could judge it. `n` is the repetition count on a published matrix, and anything below it is the signal to read: a run left out as invalid, or a metric a validator returned as `unjudged` - which shrinks that one denominator and no other. A metric that is a number or a diagnostic has no `x/n`, so it is named under the table rather than dropped from it. The table is deliberately **not** filtered by `[verdict].validity`: those metrics are columns of this very matrix, and a `delivered` column reading 10/10 by construction would hide the thing it is there to show. The **cost table** says what a run cost: tokens in, tokens out and duration, each as a median with a 95% interval from the same resampling. ```text | cell | n | in | out | duration (s) | | -------------------- | -- | ----------------------- | -------------------- | ------------ | | nothing / off | 10 | 15 929 [14 208, 17 440] | 2 286 [1 902, 2 671] | 64 [58, 79] | ``` Levels, not gaps, so they carry **no state**: an isolated measurement asserts no effect, and two intervals that do not overlap are not a result. It is computed over the runs the verdict rests on - valid, and passing `[verdict].validity` - because a run that delivered nothing is cheap by construction, and averaging it into a price makes the configuration that fails most often look like the affordable one. `n` says how many runs that left. `turns` is not here: it is a shape of the conversation rather than a price, and it is read against the reference or not at all. The **gap table** is the part a conclusion rests on. `*` established, `o` inconclusive, and no sentence may rest on an `o`. Each gap shows its `p`, Holm-adjusted over every gap of the table. When retries are present, a warning follows the table and covers even results marked established - see {doc}`../guide/invariants`. `render --reference` writes `synthesis_ref-.md` from the same measures, because a reference is a rendering choice rather than a measurement. ## `synthesis.html` Written beside the markdown, every time the markdown is - by `run`, `render` and `replay --rescore` alike. Costs no token and no network: strings in, one file out. One self-contained page, with no script, no external stylesheet and no font fetched from anywhere, because an archive is opened years later on a machine that may be offline and a page that phones home is a page that rots. It links each run's session pages once `render --html` has written them, and `render --reference` gives it the same suffix as the markdown. ## One run's directory The id is a short opaque hash, **stable** for a given scenario, cell and repetition - stable so a resume can tell an absent run from a finished one, opaque so a form can be filled without revealing the cell. Which directory it sits in is [the tree's layout](#runs-layout). `configuration.json` records what actually ran, including per-agent models and where each came from: ```json { "cell": "+subagents", "etalon": "etalon-v1", "etalon_commit": "9126095c68cbe51d7548d2b179f7470b237810df", "model": "gemma-4", "model_id": "gemma-4-31b", "thinking": "high", "injected": ["AGENTS.md", ".pi/"], "agents": { "explorer": { "model": "ilaas/gemma-4-31b", "source": "scenario override" } } } ``` Two places may declare a subagent's model, so the trace settles which one applied. :::{important} **`model` is a pattern; `model_id` is what answered.** `--model` is resolved against the models the provider actually offers, so a scenario declaring `gemma-4` runs `gemma-4-31b`. The declared value is an intention and only the session names what ran - exactly as `etalon` is a tag and `etalon_commit` is what that tag resolved to. Measured on a real matrix: six runs declared `gemma-4`, all six sessions recorded `gemma-4-31b`, and the archive said only `gemma-4`. Keeping the pattern alone left the archive unable to name the model it measured, and made a fallback to the machine's `defaultModel` indistinguishable from an ordinary resolution. `parity --smoke` now checks that what answered is still named by the pattern. `model_id` is `null` when the sessions do not say. Absence stays absence: filling the gap with the intention is precisely how a substitution would hide. ::: ### `session/*.jsonl` The agent's own trace, in the format the agent writes: **one file per attempt**, so the count matches `attempts` in `state.json`. A run that produced nothing archives its session too - it is the only evidence such a run leaves, and it is the run somebody most wants to read. Copied here byte for byte rather than left where it was written. The work directory is disposable by design, so an archive that pointed at it would keep the diff and lose the reasoning behind it on the next reboot. Relaunching an experiment **replaces** this directory, exactly as it replaces the rest. A session left by the previous launch would otherwise be attributed to this one - the file count would stop matching `attempts`, and a page rendered from the old trace would sit there looking current. :::{note} The **raw event stream is still not archived**, and the distinction matters. The *stream* is what `pi --mode json` prints while it works, almost entirely streaming deltas: 15.9 MB of it against 30 KB of session, teaching nothing the per-message record does not. The *session* is the per-message record, and that is what is kept. ::: (session-html)= ### `session/*.html`, on `render --html` Each archived session renders to a standalone page beside it, under the same stem: `session/.jsonl` gives `session/.html`. The page embeds the whole session and loads nothing from the network, so it opens from a published archive on a machine that has none. The rendering is done by `pi --export`, which is to say by the agent itself. A renderer written here would drift from the format it renders, silently, and the format is the agent's rather than ours. It costs no tokens and it is opt-in - see {doc}`cli`. ## Where clones and sessions live Under `workdir` from the config, by default `$TMPDIR/trysquare`. Deliberately outside the output directory, and deliberately disposable: the OS may purge it, and nothing of value is lost because `replay` rebuilds a tree from a tag and a diff. ```text / ├── sources/my-repo-3f2a1b9c-etalon-v1/ a [repos] URL, pinned once at the tag ├── harness/subagent-v1.2/ a [harness] brick, pinned once at its tag └── // one clone, session and trace per run ``` `sources/` exists only for repository entries that are URLs; a `[repos]` entry naming a directory is read where it already is. Because it lives under a disposable `workdir`, a `--resume` against a URL after the directory has been purged clones again, and so needs the network again. Sessions are **written** there and **copied** into the archive, which is why `parity --smoke` still takes a `--workdir`: it reads them where they were written. A run's session directory survives from one launch to the next, since the run id is stable and so the path is. So an archive takes only what the launch it belongs to produced; copying whatever happened to be there would mix a previous measurement's traces into an archive whose `measures.json` does not describe them. (naming-gap)= ## A known gap in the naming scheme The directory name carries the scenario name, etalon, provider, model and repetition count. It does **not** carry the task, the cell definitions, or the rubric. So editing a prompt - which is unquestionably part of what is measured - produces results sharing a name with the previous ones while not being comparable to them. That contradicts the principle the naming scheme exists to serve: *anything that changes what is measured changes the name.* This was found the hard way. A scenario's task had to be rewritten between two passes, because the first version enumerated its own rubric and so measured instruction-following rather than analysis. Only the repetition count kept the two sets of results apart. :::{warning} Until it is fixed: when you change a task, a cell definition or a rubric, **delete the old output directory** rather than letting a later run overwrite part of it. Two half-matrices under one name are worse than one missing matrix. A fix would add a short digest of the resolved experiment - the brick contents that actually reach the agent, plus the cells, validators and verdict - to the directory name, so that a prompt edit lands somewhere new on its own. ::: (not-implemented)= ## Not implemented yet Stated rather than left to be discovered. - **A judge is not re-scored.** `replay --rescore` re-runs script validators and reuses the archived judge verdict, because re-running a judge costs tokens. Re-scoring a judged metric therefore means measuring again. Parity layer 2 inherits the limit: it re-scores what a script can score, and names the judged metrics as out of scope.