measure/report

Source: src/measure/report.ts

What an experiment leaves behind: one JSON document, one comparison table.

Kept apart from the runner because they answer different questions - one drives the matrix, the other reads it. Two rules govern the arithmetic here.

Sums first. A summary stores totals; a mean is computed at the moment it is displayed, and never stored. Averaging averages is how a study starts lying about itself.

A failed cell stays in the report. It ran, it cost tokens, and dropping it would quietly turn “two models out of three answered” into a clean comparison of the survivors.

ExperimentModelSummary

type

export type ExperimentModelSummary = {
	/** The model pattern, as it was asked for. */
	model: string;
	/** Cells that ran. Lower than `repetitions` when the experiment was aborted. */
	runs: number;
	/** Cells whose callback reported success. */
	ok: number;
	/** Cells that failed. Counted apart: `3 runs` hides a crash. */
	failed: number;
	/** Per outcome flag, how many cells reported each value. `ok` and `error` excluded. */
	flags: Record<string, Record<string, number>>;
	/** Summed usage over the cells. Its `wallMs` is the cells' summed wall time; a mean is a display derivative. */
	total: Usage;
};

Everything one model did, across its repetitions.

ExperimentOutcome

type

export type ExperimentOutcome = {
	/** Did this cell do the work? The one flag with a column of its own. */
	ok: boolean;
	/** What went wrong. Never a column: distinct sentences compare nothing. */
	error?: string;
} & Record<string, string | number | boolean | undefined>;

What a cell reports back: a verdict, plus the flat flags the study compares.

converged, approved, iterations, rounds - whatever the callback chooses. They become the table’s columns, so they are scalars: anything the comparison needs is put here, anything else stays in the cell’s transcripts.

ExperimentReport

type

export type ExperimentReport = {
	/** What this study was called, when it was given a name. */
	name?: string;
	/** When the report was written, ISO 8601. */
	generatedAt: string;
	/** The experiment directory, absolute - the cells' `dir` is relative to it. */
	dir: string;
	/** The models compared, in the order they were given. */
	models: string[];
	/** Repetitions asked for per model, whatever was reached. */
	repetitions: number;
	/** Wall time of the experiment itself, not the sum of the cells. */
	wallMs: number;
	/** Every cell that ran, model-major. Failures included. */
	runs: ExperimentRun[];
	/** One entry per model, in the order of `models`. */
	byModel: ExperimentModelSummary[];
	/** Set when the matrix did not run whole - `"aborted"`, today. */
	error?: string;
};

The whole experiment.json document.

ExperimentRun

type

export type ExperimentRun = {
	/** The model pattern every subagent of this cell ran on. */
	model: string;
	/** 1-based, and the same number as the `rep-<n>/` directory. */
	repetition: number;
	/** The cell's directory, relative to the experiment's own - a report survives a move. */
	dir: string;
	/** The callback's verdict, promoted so the table can be read without opening `outcome`. */
	ok: boolean;
	/** Why it failed, when it did. A callback that threw lands here too. */
	error?: string;
	/** The callback's outcome, verbatim: the flags this study compares. */
	outcome: ExperimentOutcome;
	/** Wall time of this cell alone, measured around the callback. */
	wallMs: number;
	/** The cell's `usage.json` total - what the whole workflow spent. */
	usage: UsageTotal;
};

One cell of the matrix: one model, one repetition, once it has run.

experimentTable

function

export function experimentTable(report: ExperimentReport): string[] { /* … */ }

The comparison table, as Markdown lines.

Flag columns are the union of the outcome keys actually seen - a study comparing converged gets a converged column without configuring one.