Experiments¶
One workflow, M models, N repetitions, one table. This is the layer the
model knob exists for: comparing models on the same work
requires the model to be an argument, and comparing them honestly requires the
run to be repeated.
import { experiment, experimentTable, loop } from "@ai-for-dev/combo";
const report = await experiment({
models: ["ilaas/gemma-4-31b", "ilaas/gpt-oss-120b"],
repetitions: 2,
run: async (cell) => {
const result = await loop({ ...cell.options, steps: [coder, reviewer], input, until: lgtm });
return { ok: result.ok, converged: result.converged, iterations: result.iterations };
},
});
console.log(experimentTable(report).join("\n"));
That is the shape of examples/12-experiment.ts, and this is
the table it printed on those two models:
| model | runs | ok | converged | iterations | usage | mean wall | mean $ |
| --- | --- | --- | --- | --- | --- | --- | --- |
| ilaas/gemma-4-31b | 2 | 2/2 | 2/2 | 1×2 | 4 turns 368.6s ↑30k ↓21k | 184.3s | not reported |
| ilaas/gpt-oss-120b | 2 | 2/2 | 2/2 | 1×2 | 4 turns 47.9s ↑53k ↓5.8k | 24.0s | not reported |
1×2 is two runs of one iteration each. The provider reports no cost, so the
usage has no $ figure and mean $ says not reported: a missing figure is
not a zero.
An experiment is a function, not a combinator¶
It returns no Result and composes with nothing. It is a harness placed above
a workflow, and one that could be nested inside a workflow would be measuring
itself. Everything else in this library is a combinator precisely because it can
be nested; this one is deliberately not.
Which also means a flow needs no special support: runFlow takes the same
model, signal, timeoutMs, spawn and onEvent, and the cell’s directory
is its run directory.
run: async (cell) => {
const result = await runFlow(checked, input, { ...cell.options, runDir: cell.dir });
return { ok: result.ok };
},
The contract: spread cell.options¶
type ExperimentCell = {
model: string; // this cell's model
repetition: number; // 1-based, matches rep-<n>/ on disk
dir: string; // this cell's directory, absolute, already created
options: WorkflowOptions;
};
cell.options carries the cell’s model and exportDir, the experiment’s
signal, timeoutMs, cwd and spawn, and an onEvent combining the cell’s
private collector with your own listener. Spreading it is the contract: a
callback that rebuilds those by hand puts its subagents on the wrong model, in
the wrong directory, and measures nothing.
What the callback returns becomes the table’s columns:
type ExperimentOutcome = { ok: boolean; error?: string }
& Record<string, string | number | boolean | undefined>;
Flag columns are the union of the outcome keys actually seen - converged,
approved, rounds, whatever this study compares - so there is nothing to
configure. ok has its own column and error never becomes one: a column of
distinct sentences compares nothing.
On disk¶
runs/2026-09-24_00-45-16/
├── experiment.json machine-readable, every cell
├── experiment.md the table, plus the failures named under it
├── ilaas-gemma-4-31b/
│ ├── rep-1/ pi's transcripts per subagent
│ │ ├── usage.json time and tokens, attributed
│ │ └── events.jsonl the cell's whole event stream, in order
│ └── rep-2/
└── ilaas-gpt-oss-120b/
└── …
Each cell writes the same usage.json a single run writes, from
its own collector. Measurement is reused, never reinvented.
events.jsonl is the record reporter, wired
into every cell with no way to turn it off: a cell whose stream was not kept can
only be re-run, and a matrix is expensive.
The rules¶
Sequential by default.
concurrencydefaults to 1: two cells racing for the same machine measure the contention, not the models. Raise it when the providers are remote and the wall time matters more than the precision.Model-major order. Every repetition of the first model, then the second - so a matrix interrupted halfway holds finished models rather than a fragment of each.
A failed cell stays in the report, with its usage: it spent tokens before it broke, and dropping it would quietly turn “two models out of three answered” into a clean comparison of the survivors. A callback that throws is a failed cell too, not a crashed experiment.
Sums are stored, means are displayed.
experiment.jsoncarries totals; the mean wall and mean cost are computed when the table is rendered, never written down. Averaging averages is how a study starts lying about itself.An abort stops launching new cells and the partial report is still written, with
error: "aborted"at the top.
Running one¶
node examples/12-experiment.ts <modelA> <modelB>
Two repetitions of the same loop per model, the table printed at the end. Several
providers report no cost, and some report no tokens either - see
Measurements - so a usage with no $ figure, and a
mean $ of not reported, mean what they say, never “free”. Give the matrix a timeoutMs it can live with: a cell lost to a
turn that would not end is a cell missing from the comparison.
Reference¶
measure/experiment-experiment,ExperimentCell,ExperimentOptions.measure/report- the report, the table, the writes.