Experiments

One workflow, M models, N repetitions, one table. This is the layer the model knob exists for: comparing models on the same work requires the model to be an argument, and comparing them honestly requires the run to be repeated.

import { experiment, experimentTable, loop } from "@ai-for-dev/combo";

const report = await experiment({
	models: ["ilaas/gemma-4-31b", "ilaas/gpt-oss-120b"],
	repetitions: 2,
	run: async (cell) => {
		const result = await loop({ ...cell.options, steps: [coder, reviewer], input, until: lgtm });
		return { ok: result.ok, converged: result.converged, iterations: result.iterations };
	},
});

console.log(experimentTable(report).join("\n"));

That is the shape of examples/12-experiment.ts, and this is the table it printed on those two models:

| model | runs | ok | converged | iterations | usage | mean wall | mean $ |
| --- | --- | --- | --- | --- | --- | --- | --- |
| ilaas/gemma-4-31b | 2 | 2/2 | 2/2 | 1×2 | 4 turns 368.6s ↑30k ↓21k | 184.3s | not reported |
| ilaas/gpt-oss-120b | 2 | 2/2 | 2/2 | 1×2 | 4 turns 47.9s ↑53k ↓5.8k | 24.0s | not reported |

1×2 is two runs of one iteration each. The provider reports no cost, so the usage has no $ figure and mean $ says not reported: a missing figure is not a zero.

An experiment is a function, not a combinator

It returns no Result and composes with nothing. It is a harness placed above a workflow, and one that could be nested inside a workflow would be measuring itself. Everything else in this library is a combinator precisely because it can be nested; this one is deliberately not.

Which also means a flow needs no special support: runFlow takes the same model, signal, timeoutMs, spawn and onEvent, and the cell’s directory is its run directory.

run: async (cell) => {
	const result = await runFlow(checked, input, { ...cell.options, runDir: cell.dir });
	return { ok: result.ok };
},

The contract: spread cell.options

type ExperimentCell = {
	model: string;        // this cell's model
	repetition: number;   // 1-based, matches rep-<n>/ on disk
	dir: string;          // this cell's directory, absolute, already created
	options: WorkflowOptions;
};

cell.options carries the cell’s model and exportDir, the experiment’s signal, timeoutMs, cwd and spawn, and an onEvent combining the cell’s private collector with your own listener. Spreading it is the contract: a callback that rebuilds those by hand puts its subagents on the wrong model, in the wrong directory, and measures nothing.

What the callback returns becomes the table’s columns:

type ExperimentOutcome = { ok: boolean; error?: string }
	& Record<string, string | number | boolean | undefined>;

Flag columns are the union of the outcome keys actually seen - converged, approved, rounds, whatever this study compares - so there is nothing to configure. ok has its own column and error never becomes one: a column of distinct sentences compares nothing.

On disk

runs/2026-09-24_00-45-16/
├── experiment.json                machine-readable, every cell
├── experiment.md                  the table, plus the failures named under it
├── ilaas-gemma-4-31b/
│   ├── rep-1/                     pi's transcripts per subagent
│   │   ├── usage.json             time and tokens, attributed
│   │   └── events.jsonl           the cell's whole event stream, in order
│   └── rep-2/
└── ilaas-gpt-oss-120b/
    └── …

Each cell writes the same usage.json a single run writes, from its own collector. Measurement is reused, never reinvented.

events.jsonl is the record reporter, wired into every cell with no way to turn it off: a cell whose stream was not kept can only be re-run, and a matrix is expensive.

The rules

  • Sequential by default. concurrency defaults to 1: two cells racing for the same machine measure the contention, not the models. Raise it when the providers are remote and the wall time matters more than the precision.

  • Model-major order. Every repetition of the first model, then the second - so a matrix interrupted halfway holds finished models rather than a fragment of each.

  • A failed cell stays in the report, with its usage: it spent tokens before it broke, and dropping it would quietly turn “two models out of three answered” into a clean comparison of the survivors. A callback that throws is a failed cell too, not a crashed experiment.

  • Sums are stored, means are displayed. experiment.json carries totals; the mean wall and mean cost are computed when the table is rendered, never written down. Averaging averages is how a study starts lying about itself.

  • An abort stops launching new cells and the partial report is still written, with error: "aborted" at the top.

Running one

node examples/12-experiment.ts <modelA> <modelB>

Two repetitions of the same loop per model, the table printed at the end. Several providers report no cost, and some report no tokens either - see Measurements - so a usage with no $ figure, and a mean $ of not reported, mean what they say, never “free”. Give the matrix a timeoutMs it can live with: a cell lost to a turn that would not end is a cell missing from the comparison.

Reference