Display¶
Display is an observer, never a participant. No workflow may depend on a UI being present: unplug every reporter and the result is identical.
await fanOut({ agent: scout, tasks, onEvent: (event) => console.log(event) });
A reporter that throws is swallowed. A broken observer must never take a workflow down with it.
The event stream¶
The core emits, reporters subscribe and only read.
type SubagentEvent =
| { type: "spawn"; id: string; agent: string; lifetime: Lifetime }
| { type: "status"; id: string; status: "working" | "idle" | "blocked" | "done"; task?: string }
| { type: "text"; id: string; delta: string }
| { type: "tool"; id: string; name: string; args: unknown }
| { type: "post"; id: string; post: Post }
| { type: "read"; id: string; posts: readonly string[]; waiting: number }
| { type: "claim"; id: string; key: string; action: "take" | "release"; ok: boolean; heldBy?: string }
| { type: "steer"; id: string; text: string }
| { type: "usage"; id: string; usage: Usage }
| { type: "close"; id: string; result: Result }
| { type: "visit_start"; path: string; node: string; kind: string }
| { type: "visit_end"; path: string; ok: boolean; wallMs: number; usage: Usage /* ... */ };
The two visit_ events are a flow’s, and carry no
subagent id: a reader that follows subagents skips them with isVisit(event).
The task rides on the "working" transition, not on spawn: at spawn time
nobody knows yet what the subagent will be asked, and a persistent subagent is
asked several different things over its life. A reporter has no other way to
learn it, and until it did, every collapsed row in the TUI showed a blank task.
onEvent takes a single listener, so watching in two places at once means
composing:
import { combineReporters, createHerdrReporter, createRunPicture } from "@ai-for-dev/combo";
const picture = createRunPicture();
onEvent: combineReporters(picture.reporter, createHerdrReporter());
createHerdrReporter() returns undefined outside herdr, and combineReporters
drops it.
Reading is an event too, and a swarm cannot be read back
without it. The posts say who said what; read says who knew what, and a
member handed nothing is the strongest thing the record holds about what a
member could not have known. Three members claiming the same file reads as three
models thinking alike until the three reads before them are in the stream.
Picking a reporter¶
import { autoReporter, consoleReporter, silentReporter } from "@ai-for-dev/combo";
onEvent: autoReporter(); // herdr when it is running, silent otherwise
autoReporter() never warns and never throws: not running under herdr is the
normal case, not a degraded one.
consoleReporter() prints one line per event, board traffic included:
⇣ member#3 was handed nothing
⚑ member#3 take console.ts → granted
⚑ member#2 take console.ts → refused (member#3)
✉ member#2 → member#3 [ask] Are you far off on console.ts?
⚑ member#3 release console.ts → given back
A refusal names the holder: contention has somebody in it, and “refused” alone cannot tell that from asking for something that was never there. The TUI widget shows subagents rather than traffic, and says nothing about a board.
Keeping the stream¶
import { combineReporters, recordReporter } from "@ai-for-dev/combo";
onEvent: combineReporters(autoReporter(), recordReporter("runs/latest/events.jsonl"));
One JSON object per line, {"ts": <ISO 8601>, …event}, in the order things
happened. pi’s own JSONL already holds each subagent’s transcript; what this adds
is the two things that live between them - the interleaving (who was
working while who else was reading) and our timestamps, since pi has no
notion of the wall clock a workflow runs on.
Events are recorded verbatim, with no filtering: a recorder that edits its own
record is worse than a large file, and the analysis it exists for is the one
nobody planned in advance. It writes with appendFileSync - one syscall per
event, deliberately, because the run worth reading afterwards is the interrupted
one and a buffered stream loses its tail exactly then.
Every experiment cell gets one, at events.jsonl next to its
usage.json. There is no option to turn that off.
herdr¶
Inside herdr, a subagent can get its own split and show what it is doing:
await fanOut({
agent: scout,
tasks,
concurrency: 3,
openInHerdr: true, // opt-in, per subagent
onEvent: autoReporter(),
});
A herdr pane cannot host an in-process subagent: there is no process and no TTY
to attach. So the pane hosts a client of the mirror
instead: it runs pane/main.ts, which attaches to the subagent by id, draws
the session with pi’s own components and sends the keyboard back. What you see
in the split is the task, the answer as Markdown and each tool call in its box,
the way pi shows its own session, and what you type into it reaches the
subagent while it works. Splits close on their own when their subagent does,
so a fan-out leaves no orphan panes.
Opening one takes three calls, because herdr has no single call that does all
three: pane.split makes the pane beside ours and answers with its id,
pane.rename puts the subagent’s name on it (for a flow’s subagent, its
agent and where the flow keeps it: coder @ deliver#2/work[1]/pair), and
pane.send_input types the
client’s command into the shell the split started. agent.start sounds like the call
that opens one and is not: it puts a recognised agent into a pane that already
exists, and mistaking the two is how /herdr on once opened nothing at all
(decisions). node scripts/check-herdr.ts holds every one of
those calls to herdr api schema --json.
Detection needs HERDR_ENV=1, HERDR_SOCKET_PATH and HERDR_PANE_ID. All
three, or nothing at all. The pane id is what the splits open beside, so they
land next to ours rather than next to whichever pane another client has
focused.
openInHerdr is opt-in per subagent, exactly like lifetime, so a fan-out of
twenty branches cannot carpet the screen by accident. It can be a default on the
agent itself, which is often what you want:
---
name: scout
description: Locates the code relevant to a question
tools: read, grep, find, ls
openInHerdr: true
---
The other regime - watch everything - belongs to the reporter, not to the core, because who gets a pane is a display decision and the workflow runs identically either way:
/herdr on # for this pi session
createHerdrReporter({ all: true }); // from a script
There is no environment variable for it, and there will not be: configuration is an argument or a command, never something a shell exported three days ago.
/herdr on asks herdr before answering, and says what it heard:
herdr: every subagent gets its own split
herdr: on, but pi is not running inside herdr - nothing will open
herdr: on, but herdr refused pane.split: missing field `kind` - nothing will open
The third line is the one worth having. A herdr that is there and refuses every
request looks exactly like a herdr that works, and the reporter watching it may
not warn about anything at all. The command may, so it is where a refusal gets
said. The probe sends the same pane.split a run sends, with a
target_pane_id no herdr can have: pane_not_found means the request was
understood, anything else is the message above, and neither opens a pane. The
preference is set either way, because what you asked for is not herdr’s to
decide.
The board gets a pane of its own¶
A swarm is the case a pane per member does not cover: what one
member said is in its own pane, and who it was talking to is only legible where
all of them are. So the first thing anybody says opens one more pane, named
board, carrying the exchange in order:
⇣ member#2 was handed nothing
✉ member#1 → member#2 [ask] who has console.ts?
⚑ member#2 take console.ts → refused (member#1)
Everything on it is announced by the board and the claims themselves, whoever
acted on them - a member through its tool, or the swarm handing a member what it
has not seen and taking back what a member that stopped was holding - so the
pane, the console and events.jsonl hold the whole of the traffic and not the
half that went through a tool. Each member’s own pane shows its half of that as
the tool calls it made, the way pi shows a tool call. The board’s wording is the console reporter’s, from
one place (src/reporters/traffic.ts), because two displays of one run are
read side by side and a difference between them would read as a difference in
the run. The board is the one pane that follows a file rather than the mirror:
nobody works in it, so there is nothing to type to.
The pane opens only when the run is being watched at all, closes with the last member, and carries no herdr agent: nobody works in it, so there is nothing to report a state for and nothing to release.
Reaching in: the mirror¶
Every live subagent is registered, by id, with the mirror
(src/mirror.ts): a unix socket server, one per process, that starts the
first time something asks where it is and goes with the last subagent. A
client attaches by id and speaks newline-delimited JSON:
→ { "attach": "scout#1" }
← { "type": "attached", "id": "scout#1", "agent": "scout", "model": "…", "cwd": "…", "pi": "file:///…/pi-coding-agent/dist/index.js" }
← { "type": "message", "message": … } // the transcript so far, one per entry
← { "type": "message_start", … } // then pi's own events, as pi emits them
← { "type": "status", "id": "scout#1", … } // and ours, around them
→ { "type": "steer", "text": "look at test/ first" }
→ { "type": "abort" }
← { "type": "close", … } // then the line goes
attached names the pi package this process runs, because the events that
follow are that pi’s and whatever draws them has to be built from the same
one. A message_update arrives per token and carries the whole partial
message, so the mirror sends the latest at most every 50 ms, and always before
whatever follows it.
A steer is delivered only while the subagent works. Measured: a steer
queued on an idle session is delivered with the next prompt() and answered
in place of it, and a follow-up is answered inside the next prompt()
after the task, so the workflow reads a person’s exchange back as its own
result. Between tasks the mirror answers
{ "type": "refused", "reason": "between tasks - it can only be steered while it works" }
and queues nothing. abort is stop(), safe at any moment, and the turn comes
back stopped like /stop.
Every steer that went through is also a steer event on the run’s stream, so
the record and the console (⌨ scout#1 ← look at test/ first) hold that a
person spoke; the pane it was typed in draws it as pi draws a user message,
once the subagent takes it. A run somebody steered is not the run they would
have got by watching, and two identical events.jsonl must not describe two
different runs.
This is not a reporter. A reporter reads the stream and never reaches the
session; the mirror is a port of the core, beside ask and verify, and it is
opened by the same act that opens a pane. Unplug every reporter, attach nobody,
and the result is identical.
The pane¶
pane/main.ts is the client a split runs:
node pane/main.ts --socket <mirror socket> --id scout#1
It draws the session with pi’s own components, so it reads like pi: the task
in a user box, the answer as Markdown, each tool call in its box with the
result folded under it, and a prompt at the bottom. Two dim lines under the
prompt say who this is and what it spent (scout#1 · provider/model · 1 turn 5.7s ↑15k ↓385), then what it is doing and what the keys do:
Key |
While it works |
Once it is done |
|---|---|---|
|
steers: the line reaches the subagent after the tool call in flight |
|
|
stops it, like |
closes the pane |
|
leaves the pane; the subagent does not notice |
the same |
A line typed between tasks is answered under the prompt with the mirror’s refusal, and nothing is queued. A closed subagent leaves its transcript on screen, so the last thing it said is not lost with the window.
The pane’s theme is the user’s, read the way pi reads it. That is not the subagent inheriting anything: the pane is the user’s window, and its colours are theirs.
The pi TUI¶
While the subagents work, a dot per subagent sits just above the prompt:
● scout#1 grep /lifetime/
provider/model · ↑12k ↓209 · 12.4s
✓ scout#2 provider/model · ↑8k ↓150 · 8.1s
● while it works, ✓ when it succeeded, ✗ when it failed, coloured by
status. That reading is made once, by standingOf, and the tool’s card, the
summary table and the console draw the glyph it gives them rather than deciding
their own; the dimmed line underneath carries model, tokens and elapsed time,
counting up live. Events alone cannot keep that clock - usage.busyMs only
lands when a turn ends - so the widget reads the turn’s start and repaints on a
timer. A subagent thinking for twenty seconds emits nothing, and a frozen clock
reads as a hung agent.
The tokens wait for a turn to end, because that is when pi’s counters are
read. A subagent in its first turn shows its model and its clock alone,
provider/model · 3.1s, and ↑0 ↓0 there would be a figure nobody measured.
After a turn, a provider that reports no tokens reads ↑0 ↓0: that zero was
read. The cost follows the same rule: it is shown when pi reported one, and
left out when it did not, since $0.0000 reads as free. The tool row’s totals
and the summary table do the same, and an experiment’s mean $ column says
not reported.
A tool call that came back an error, refused because the agent has no such
tool or failed while it ran, is marked ✗ with the first line pi said about
it, in the activity and in the tool row alike:
✗ write notes.txt · Tool write not found. pi’s event does not say which of
the two it was, so the words are pi’s.
Two lines while it works, one once it is over, as scout#2 above. The
second line of a finished subagent held its last tool call, which nobody needs
any more, so its numbers move up beside the tick and the line goes. A fan-out of
three takes seven lines at its widest and shrinks as it finishes, rather than
holding the terminal at its widest until the run ends. A failure keeps what the
tick cannot say: ✗ coder#1 402 from the provider provider/model · 3.1s.
A subagent that was delegated sits under the one that asked for it, here and in the tool row alike:
● explorer#1 subagent explorer → scout
provider/model · ↑8k ↓412 · 21.0s
● scout#1 read src/reporters/tui.ts
provider/model · ↑14k ↓980 · 9.4s
A row says how deep it sits and the drawing applies the indent, which is the same split as everywhere else here: the picture knows the depth, the terminal decides what a level looks like. A run with no delegation is drawn exactly as it always was.
A run that walks a flow draws the flow’s plan instead, filled as its
visits go: a running visit expanded, with each of its subagents under it on one
line, what it is doing, its model, tokens and clock; an ended one a single
line; what is not visited yet dimmed under ○. The plan takes sixteen rows at
most, cut above and below what runs now, and it is a component painted at the
width pi gives it, since pi cuts a widget given as lines at ten. Every line is
cut to that width in terminal columns, colour codes and wide characters
counted as the terminal counts them: pi stops on a line one column too wide.
The summary line’s time is how long the run has gone on, by the view’s clock,
and each ended line counts what every life of the run spent there. See
Extension.
Stopping what you are watching¶
The dots are also the list of what can be called off. Under them, while something is still running:
esc stops everything · ctrl+↑↓ selects · ctrl+del stops the selected one
esc stops every subagent of the run. It is not intercepted: pi’s own
interrupt fires as well, so inside a model’s turn the turn goes with the
subagents and the model gets no chance to delegate again. During /run or
/step pi has no turn to abort, and this is what stops them.
The one time it does not is while a question card is up. There esc is the
card’s, and means what its help line says. On the interview’s card that is
“write the brief with what you have”: the interviewer that will write the brief
is a subagent of the same run, and a key that stopped it as well would end the
interview it was meant to close. On a flow’s card that offers no “enough”, it
is the run’s stop, pressed on the card that holds it. The run’s subagents are idle
while a question waits, so nothing is running that the key would have called
off.
ctrl+↑ and ctrl+↓ move a ▸ through the subagents that are still working -
delegated children included, in the order the widget draws them - and ctrl+del
stops the one it points at. /stop does the same by name:
Command |
What it stops |
|---|---|
|
the selected subagent, or the only one running |
|
that one, e.g. |
|
the whole run, like |
One subagent stopping is not the run stopping: its turn comes back as a
failed Result reading stopped, and what the workflow does next is the
workflow’s business - a fanOut branch dies alone, while a flow’s node that
fails ends the run unless its on-fail: says to go on. Either way what already
ran is kept, and exported.
/stop is only typeable while the model is running subagents: pi executes an
extension command immediately during a turn, but processes no submission at all
while a slash command of its own is awaiting. That is why the same act has a
key.
The widget disappears the moment the work ends, in a finally, so a thrown
workflow never leaves a dead row of dots above the prompt. The full record is one
line below, in the tool row: one line per subagent with its last tool calls.
Expand it - the hint comes from your own keybinding configuration, not a
hard-coded Ctrl+O - for the full task, every tool call, the output rendered as
Markdown, and usage per subagent.
A parallel run shows what it achieved (2/3 done, 1 running), and a loop says
whether it converged or merely ran out of iterations.
One picture, many readers¶
The event stream is folded once, by createRunPicture(), into the picture of
the run: every subagent with its status, task, tool calls, usage and how deep
it sits under the one that delegated to it, plus what they add up to. The TUI
widget, the console reporter, usage.json, an experiment’s cell and the
subagent tool’s own result all read that picture. None of them folds the
stream itself, so a fix to how the picture is built - the launch order of a
fan-out, the depth of a delegated subagent - reaches every one of them.
snapshotFrom(subagents) derives the counts and the total from a list of
subagents alone. It is what the live picture calls on every frame, and what the
extension calls on the subagents pi handed back serialised in a tool result, so
both ends draw from the same arithmetic.
Formatting sits apart, in tui.ts: it takes a snapshot in and gives strings
and rows back, with no pi-tui import. The extension draws them. A line is
therefore tested by calling the function that makes it, never by scraping a
terminal, and the same picture would feed a web view without touching a
component.
A flow run has a second fold beside the picture: livePlan in src/flow/
reads the visit events and the journal into the flow’s plan, and ignores
the subagent events except spawn, whose visit ties a subagent to the
plan line it works for. The picture skips the visit events, so each fold
reads its own half of one stream and neither is changed for the other.
The live view says how the plan fills.
Reference¶
events-SubagentEvent,EventBus.reporters/index-autoReporter,combineReporters.mirror-registerMirror,mirrorSocket, the wire.reporters/picture-createRunPicture,snapshotFrom,RunSnapshot.reporters/tree-treeOrder, a child under its parent.reporters/tui-widgetRows,summaryTable, the formatting.reporters/record-recordReporter, the event stream on disk.