This is the walkthrough for a consumer of processkit-cli — an orchestrator or
adapter (in particular the processkit-py CLI) that launches runs through this
binary and reads its results back — rather than for a contributor to this
repository (see docs/architecture.md for that audience). It
ties together, in the order an adapter actually exercises them, the five
normative documents that each cover one part of the compatibility surface on
their own: docs/schema.md, docs/exit-codes.md,
docs/control-plane.md, docs/registry.md,
and docs/compatibility.md. This document does not restate
their normative text — every concrete claim below is a pointer to, and a minimal
worked example of, the contract those documents define; on any disagreement,
the linked document is the source of truth.
Before launching anything through a candidate processkit-cli binary, verify
it is compatible. probe is side-effect-free — it spawns no child and touches
no registry or container — and prints one JSON report line to stdout:
processkit-cli probe --json \
--require-schema-version 1 \
--require-exit-code-band 100-119 \
--require-surface run:--jsonl \
--require-surface run:--capture-dir \
--require-surface inspect:--json \
--require-surface cancel:--run-id \
--require-surface kill:--run-idThe report (one line, shown reformatted here):
{
"probe_version": 1,
"binary": "processkit-cli",
"version": "0.2.2",
"schema_version": 1,
"exit_code_band": { "start": 100, "end": 119 },
"surface": ["cancel", "cancel:--run-id", "inspect", "..."],
"compatible": true,
"mismatches": []
}- Pin
schema_version(--require-schema-version <N>) and the reserved exit-code band (--require-exit-code-band <start>-<end>) so a future breaking change is caught here, before a run, rather than by a JSONL parser or an exit-code table drifting silently out of sync. - Pin the exact CLI flags the adapter is about to use with one
--require-surface <token>per token (a bare subcommand name, or<subcommand>:--<long-flag>) — this is how an adapter confirms a flag it depends on (for examplerun:--capture-dir) actually exists on this build before passing it. - Pin
run:resource-summaryif the adapter will read what a run consumed — peak memory, CPU, IO bytes, peak process count. Like the token below it this is a capability, not a spelling (note the missing--), and it says this build'srunemits the terminalresource_summaryevent. Requiring it is the only way to learn that before a run: an older binary simply writes no such line, which an adapter would otherwise discover by finding it missing from a stream it has already finished collecting. The token guarantees the event, not any particular number in it — which measurements are populated follows the containment mechanism and, for memory and CPU on Linux cgroup v2, how the run ended, perdocs/resource-limits.md. Read that matrix before building an adapter around a specific axis; pinning the token is not the same as learning a number will be there. - Pin
attest:peer-identityif the adapter will gate anything on containment membership (§4). It is the other capability rather than a spelling — note the missing--— and it says this build can obtain a kernel-authenticated identity for a control-plane client on this platform, which is what makesattestable to answer at all. Requiring it turns "this platform cannot prove membership" into an ordinary fail-closedPROBE_INCOMPATIBLE(110) here, instead of apeer_identity_unsupportedrefusal in the middle of a job. Its presence is a guarantee; its absence withholds one rather than predicting failure, so an adapter that requires it is choosing not to depend on an unguaranteed capability. - An unmet expectation makes
probeexitPROBE_INCOMPATIBLE(110) withcompatible: falseand the concretemismatches; a malformed--require-*argument (not an incompatibility, a bad flag) is the ordinaryUSAGE(100). A satisfied — or unrequested — surface exits0.
This is a fail-closed contract: an adapter that skips the preflight (or
silently proceeds after a PROBE_INCOMPATIBLE) re-introduces exactly the
uncontained-launch hazard this project exists to prevent. See
src/probe.rs and the normative exit-code table in
docs/exit-codes.md.
The report's shape is published as a JSON Schema with a golden fixture —
fixtures/schema/cli/probe.schema.json and probe.jsonl — so an adapter can
validate what it parsed instead of re-deriving the shape by hand. Every
machine-readable output in this guide has such a pair; see
fixtures/schema/cli/README.md
for the full table, and docs/compatibility.md, "Machine-output schemas", for
why some of these outputs carry no version field of their own (the probe
report's probe_version, the inspect snapshot's snapshot_version, the
failure envelope's error_version — §7 — the attestation's
attestation_version — §4 — and the qualification report's doctor_version
— below — are the five that do).
probe --print-schema is a separate, simpler mode on the same subcommand: it
prints this binary's embedded JSONL event-schema document instead of the
report above and exits 0, so an adapter that only needs the schema for its
own version — no clone, no tag to match — can fetch it offline without a
compatibility check. It cannot be combined with any --require-* flag:
that combination is rejected as an ordinary USAGE (100) parse error, never a
silent skip of the requested checks, so it can never produce a false "ok" on
an invocation that also asked probe to verify expectations. See
docs/schema.md, "Getting the schema without a git checkout".
A passing probe says the binary you found is the one you need. It does not say this
machine can run a contained process — by construction, since proving that would mean
running one, and a preflight that spawned a child would not be a preflight. The two
claims come apart in practice: a registry directory that cannot be created or is not
owner-only, a containment mechanism the kernel will not hand out, a local IPC endpoint
that will not bind. Each of those passes every --require-* check above and fails the
first production run.
doctor closes that gap by doing the thing:
processkit-cli doctor --json --require-abrupt-cleanup whole_tree --check-resource-controller --require-resource-controllerIt performs a bounded scratch run of this binary's own harmless child
(doctor --scratch-child), drives that run through the ordinary control plane
(inspect, cancel, terminal wait), confirms teardown left nothing, and reports the
facts it observed — the registry directory and its owner-only protection, the
containment mechanism and abrupt-cleanup level this host really gives a run, the
transport round-trip, a confirmed-empty cleanup, and per-phase timings. On success it
leaves nothing behind; on a failed phase it keeps a diagnostics directory and names it
in the report (diagnostics_dir).
Where it belongs in an adapter's flow: once per host, at setup or install time, not
before every run. It is the counterpart of the probe above, on the other axis —
probe |
doctor |
|
|---|---|---|
| Subject | This binary | This host |
| Side effects | None: no child, no registry, no container, no endpoint | A real scratch run: registry entry, container, control endpoint, all cleaned up |
| Cost | Milliseconds | Under a second, bounded by --timeout (default 30s) |
| Run it | Before every launch, or at least whenever the binary may have changed | Once per host, at setup time — or when a host starts behaving differently |
| Fail-closed code | PROBE_INCOMPATIBLE (110) |
HOST_UNQUALIFIED (116) |
The requirement flags gate the exit code only; the report carries the observed
facts either way, so an adapter can act on the code and still log everything the
qualification saw. --require-resource-controller needs
--check-resource-controller alongside it — a requirement about a fact this
invocation never observed is refused as a USAGE (100) rather than guessed.
doctor --json's shape is published like every other machine-readable output here:
fixtures/schema/cli/doctor.schema.json and doctor.jsonl. See
docs/troubleshooting.md, "Qualifying a host: doctor", for
reading a negative verdict, and docs/exit-codes.md, "Qualifying a
host: doctor", for why 116 is its own code.
The recommended invocation for an adapter:
processkit-cli run \
--run-id build-42 \
--jsonl .processkit/build-42.jsonl \
--capture-dir .processkit/build-42/capture \
--env-clear \
--env PATH="$PATH" \
--env-remove CI_SECRET_TOKEN \
--timeout 10m \
--grace 5s \
-- dotnet build--jsonl <file>is the only place lifecycle events are written — never stdout, so the child's own stdout/stderr stay pristine. Give every run a distinct path; the file is created or truncated at the start of the run.--run-id <id>is the identifierinspect/cancel/kill/attestlater match on — supply one you control (rather than the generated default) so the supervision step (§4) has a stable handle. Two live runs sharing one--run-idis legal but makes every supervision command against it fail closed as ambiguous (§4, §6) — keep run ids unique across an adapter's own concurrently-live runs.--capture-dir <dir>additionally tees stdout/stderr to<dir>/stdout.log/<dir>/stderr.logwith a byte count, a SHA-256, and explicit truncation/write-error flags per stream (theoutput_capturedevent, §3) — use this when the adapter needs the transcript as a file rather than (or in addition to) the live echo.--no-echosuppresses the runner's own live retransmission of the child's stdout/stderr — the exact "pure noise" an adapter reading results from--jsonl/--capture-diralone does not want interleaved with its own output. The pipe,--capture-dir, and the JSONL stream are all unaffected; it conflicts with--inherit-stdio, which runs no pump to suppress in the first place.--detachreturns as soon as the run has provably started instead of blocking for its whole duration — the "launch and let go" shape, for an adapter that supervises out of band (§4) rather than by staying the runner's parent. It re-spawns the CLI detached (a new session on Unix, aDETACHED_PROCESSon Windows) and waits only until that copy has registered the run and writtenrun_startedto--jsonl, so on return the run is already visible tolist/inspect/wait. An adapter that captures the launch command's output (subprocess.run(..., capture_output=True)) gets end-of-file when the call returns, not when the run ends: the detached runner keeps none of the caller's pipes open. The exit code changes meaning under this flag and only under it: it reports the start —0once the run started, or the same reserved code the failure would have produced in the foreground (a missing program is stillSPAWN101) — never the child's own code, which stays in the terminalrunner_exitevent (§3). Adapters that need the child's result must read it there, or viawaitplus the event stream. It conflicts with--inherit-stdio/--inherit-stdin(nothing interactive survives detaching) and implies--no-echo's discarding sinks, while--jsonl,--capture-dir,--idle-timeout, and--snapshot-intervalbehave exactly as they do in the foreground. The detached runner's own stderr isnull, so--jsonlis the only channel that reports anything — including a failed member read, which is why that failure is a flagged event rather than a warning (§3). On Windows, pair it with--create-no-windowfor a console child: the detached runner has no console to lend it, so the OS gives the child one of its own. Seedocs/exit-codes.md, "Detached runs".--env-clear/--env-remove <KEY>/--env-file <file>/--env <KEY=VALUE>give the adapter control over the child's environment, applied in that fixed order — clear, then remove, then files, then explicit sets — regardless of flag order on the command line, so an explicit--envalways wins on a duplicated key. SeeREADME.md, "Environment", for the full precedence rule.--run-id-env <KEY>hands the child the run's final id — the--run-idabove, or the generated one when the adapter did not supply an id — in the named environment variable, applied after every flag in the previous bullet. This is the alternative to minting an identity adapter-side and passing it twice (--run-id <id> --env KEY=<id>): one value, no second copy to drift, and it is the only way to give the child a runner-generated id, which is otherwise not knowable until the run has already started. It is opt-in (no key is injected by default) and it composes with--detach— the detached copy is re-spawned with an explicit--run-idfor the id its caller already reported, so the child sees the same value the caller has. An explicit--env <KEY>=…for the same key is refused as aUSAGE(100) parse error rather than silently overridden — "the same key" by the platform's own rule, so on Windows an--enventry differing from<KEY>only in case is that same refusal. Treat the value as correlation data only: it identifies a run, it does not authenticate one, and any process able to set an environment variable can forge it.--max-memory <size>/--max-processes <n>/--cpu-quota <cores>cap the run's whole process tree. Enforcement needs a real container (Windows Job Object or Linux cgroup v2 at the real hierarchy root); where the platform or environment can't apply a cap the run fails fast with alimit_hitevent (§3) andBACKEND(102) rather than running silently unbounded — so an adapter that depends on a cap must treat alimit_hitas a hard failure, not a warning. See the resource-limit platform matrix, which has a separate FreeBSD process-reaper row — normal whole-tree containment, membership, and kill/teardown, but no memory/process/CPU caps orresource_summaryaccounting — and the distinct macOS/non-FreeBSD BSD process-group row; cgroup v2 is often unenforceable under systemd/containers/typical CI, and the Linux--max-processescaveat.- Command-line redaction.
run_started'scommandfield is redacted by default: the raw argv is not recorded, only a one-way SHA-256 fingerprint (argv_sha256) and a classified worker-shapehint(both derived from argv but unable to reveal it) — filled on every run whether or not--argv-rawis given. Pass--argv-rawonly when the adapter's own storage for the resulting JSONL is at least as trusted as the command line itself; do not default to it. Seedocs/schema.md("Command redaction") for the exact fingerprint encoding, which an adapter reproducing the digest independently must match byte for byte. That encoding is defined over each element's canonical bytes, which are platform-specific for an argument that is not valid Unicode (Unix: the argument's own bytes; Windows: the WTF-8 encoding of its UTF-16 code units) — an adapter that hashes a UTF-8-decoded string instead will agree on every ordinary command line and disagree on exactly those. The same case is the one where--argv-raw'sargvarray carries an escaped element rather than the argument verbatim; an adapter that re-runs or compares a recorded argv must decode it. An adapter that stores what it read must also check its sink: an escaped element is the only string value in this schema that can contain U+0000, and while the JSONL line itself stays NUL-free text (the wire form is the ordinary six-character escape), a decoded value handed to a sink that forbids NUL is rejected or truncated — PostgreSQL refuses ajsonbdocument containing it, so an adapter ingesting whole event lines intojsonbloses the entire event rather than one field. All three rules are indocs/schema.md.
run is not shell-free by accident — everything after -- is the literal
<program> <args...>, with no shell to expand or reinterpret it; an adapter
that needs shell features passes the shell as the program explicitly.
--jsonl accumulates one JSON object per line as the run proceeds; parse it as
newline-delimited JSON, dispatching on each object's event field. A minimal
reader:
import json
with open(jsonl_path, encoding="utf-8") as f:
for line in f:
evt = json.loads(line)
if evt["schema_version"] != 1:
raise IncompatibleSchema(evt["schema_version"])
handle(evt["event"], evt)Pin schema_version here too (or rely on the probe preflight in §1 to have
already ruled out a mismatch) — never assume a fixed shape without checking it.
Or let the binary read it back for you. events is the first-party reader of
the same stream, so an adapter that only needs to show a run's story — or to
check one — does not have to write the loop above at all:
processkit-cli events --run-id build-42 # rendered for a human
processkit-cli events --run-id build-42 --follow # ... as it happens
processkit-cli events --file "$jsonl" --json # the runner's own bytes
processkit-cli events --file "$jsonl" --validate # conformance checkIt resolves the stream through the registry (--run-id, the same jsonl locator
list --json publishes in §5) or reads a path directly (--file, for a stream
whose registry record is already gone — a clean exit deletes its own record — or
one this registry never knew about). Exactly one of the two is required and they
are mutually exclusive; there is no precedence rule, so passing both is a USAGE
(100) error. Like list/wait it is read-only: registry opened read-only, no
control-plane round trip, nothing mutated.
Three properties matter for an adapter:
--jsonis a pass-through, not a re-serialization. Each line is emitted byte for byte as the runner wrote it, so a field a newer runner added survives the trip — the pipelineprocesskit-cli events --file … --json | your-parseris exactly as lossless as reading the file yourself. A line that is not JSON is reported on stderr instead of emitted, so stdout stays parseable JSONL.--followis bounded by the run, never by an invented deadline. It returns at the terminalrunner_exit, or once the registry reports the run over and the stream has stopped growing — the abrupt-death case, explained on stderr rather than passed off as a complete stream. It hands out only complete lines, so a half-written event is never parsed as an event.--validateis a conformance gate. It checks every line against the schema document this binary embeds — the same oneprobe --print-schemaprints in §1 — reports each violation by line number and by what it violated, and exitsEVENTS_INVALID(114) if any line fails,0if none does. An unreadable stream is stillSETUP(111) and a--run-idnaming no single stream is stillCONTROL(103), so a fixture-checking CI job can tell "invalid" from "could not be checked". This is the recommended way for an adapter to keep its own recorded fixtures honest against the runner version it targets.
Ordering (normative: docs/schema.md). A normal run
emits, in order:
run_started— the child was spawned; carriesrun_id,root_pid, containmentmechanism, theabrupt_cleanuptri-state, and the redactedcommand.members_snapshot(reason: "spawn") — the container's members at that point. Exactly one by default; a run started with--snapshot-interval <duration>emits additionalmembers_snapshotevents (reason: "interval") on that cadence, all of them after this one and all of them before step 3 — never inside the teardown pair. Route by event type and treat the count as open-ended: within a schema version an adapter must not assume an event type it knows occurs only once (the full list of what a reader must tolerate within a version is indocs/compatibility.md). Every one of these events carriesread_error; when it istruethe read failed andmembersis an empty fallback, not a confirmed-empty tree — check the flag before drawing a conclusion about the tree from an empty array.- Either the natural-exit path (
root_exited,cleanup_started,cleanup_finished) or a runner-imposed ending's reason event (timeout,cancelled, orkilled) followed by the samecleanup_started/cleanup_finishedpair. resource_summary— what the tree consumed, inserted between the ending event of step 3 and itscleanup_started. Exactly one, on every run that spawned the child: no flag turns it on and no platform turns it off (alimit_evidence, when a cap was requested, sits immediately before it). Every measurement is independently nullable,nullmeans "this mechanism, at this read point, does not account for it" rather than zero, andread_errormust be checked before concluding anything from anull. The event's presence is guaranteed; its contents are not — on Linux cgroup v2 an ordinary run whose child exited on its own reports all five asnullwithread_error: false. Seedocs/schema.mdand the platform matrix indocs/resource-limits.md.output_captured, only when--capture-dirwas set.runner_exit— always the last line, the terminal event of every run, including a runner failure before the child ever started (in which casespawn_failedorcontainer_failedprecedes it instead, with norun_startedand noresource_summary— there was no tree to measure).
Telling outcomes apart. Two signals distinguish how a run ended, and an
adapter should use both together: the process's own exit code (fastest to
check, no parsing needed) and the terminal runner_exit event's source and
code fields (authoritative — see docs/exit-codes.md,
"Why a band is not enough on its own"):
runner_exit.source |
Exit code | Meaning |
|---|---|---|
child_exit |
the child's own code (child_code, echoed in code too) |
The child ran to completion on its own. |
timeout |
106 |
A runner deadline elapsed and the runner tore the tree down — the whole-run --timeout or the --idle-timeout (child silent past the idle window). The preceding timeout event's reason (overall / idle) says which; both reuse this one source and code. |
output_overflow |
113 |
A capture stream crossed --capture-max-bytes while --capture-overflow cancel was active, and the runner tore the tree down through the same graceful path a timeout takes. The preceding output_overflow event names the stream that crossed first and the max_bytes ceiling (see docs/schema.md). |
cancelled |
107 |
A local stop signal cancelled the run: a Ctrl-C, on Unix a SIGTERM/SIGHUP (an external kill/systemctl stop/cancelled CI job, or a hung-up terminal), or on Windows a Ctrl-Break/console close/logoff/system shutdown. The preceding cancelled event's source (ctrl_c / sigterm / sighup / ctrl_break / ctrl_close / ctrl_logoff / ctrl_shutdown) says which; all reuse this one source and code. |
control_cancel |
108 |
A control-plane cancel (§4) cancelled the run. |
control_kill |
109 |
A control-plane kill (§4) force-killed the run. |
spawn_error |
101 |
The child never started (spawn_failed precedes it). |
container_error |
102 |
The container could not be created or joined (container_failed precedes it) — including a requested resource limit (--max-memory/--max-processes/--cpu-quota) the platform could not apply, in which case a limit_hit naming the limit precedes the container_failed (see docs/schema.md). |
internal |
104 |
A genuine runner bug — the runner's own logic hit a state it rules out. |
setup |
111 |
An ordinary fail-closed setup failure (an unwritable --jsonl/--capture-dir, an unreadable --stdin-file) — distinct from internal, and the caller can usually act on it (bad path, permissions, resources). |
Only source: "child_exit" carries a non-null child_code; every other
source means the child's own exit code was never produced or is not what
code reports, and child_code is null. See the full field reference in
docs/schema.md and the exit-code contract in
docs/exit-codes.md.
Telling the two readings of a reserved-band code apart without opening the
stream. The table above is the authoritative answer, and it costs a file read after
every call: a 106 is the runner's TIMEOUT or a child that happened to exit
106, and only runner_exit.source settles it. There is a cheaper answer for the
common case — and it is why this guide offers no terminal receipt file (§8). Under the
global --error-format json (§7), a runner-owned ending prints exactly one
envelope line on stderr and a child's own exit prints none, so the envelope's
presence — not the numeric code — separates the two readings:
processkit-cli --error-format json run --run-id build-42 \
--jsonl .processkit/build-42.jsonl --no-echo -- ./build.sh 2>build-42.err
rc=$?
# rc=106, no envelope line in build-42.err -> the child itself exited 106
# rc=106, one envelope line in it -> {"error_version":1,"code":106,"kind":"timeout",...}Test for the envelope's presence, not for an empty file. Even with the echo
suppressed, stderr may carry processkit-cli: warning: … prose lines — a registry
that could not be opened, event logging that switched itself off, a member read that
failed. Those are not envelopes and not failures, and they keep their prose in both
modes (see docs/exit-codes.md,
"What the envelope does not cover"), so an adapter that reads an empty file as "the
child's own code" will misread a genuine 106 the moment any harmless warning
appears. Look for the one line that parses as JSON carrying error_version: there is
at most one per invocation.
The envelope's kind is spelled exactly like the source column of the table above —
every row of it but child_exit, which is by definition the reading that prints no
envelope — so the two channels need no separate vocabulary. --no-echo is what keeps
the child's own bytes off the runner's stderr, and it is the only flag that does —
--capture-dir is an independent axis that files a transcript without suppressing
anything, so pass it alongside --no-echo when the child's output is still wanted
(the pump stays wired either way, so the transcript is complete). With the echo left
on, the envelope is still emitted — it is the runner's own final stderr line — but
the child's stderr is interleaved with it, which is one more reason to match on the
line rather than on the stream as a whole.
This does not make the stream optional, and it is not a second outcome artifact.
--jsonl is required on every run, the envelope reports only the runner-owned
endings, and everything else a supervisor reads — the containment mechanism and
abrupt_cleanup level, root_pid, the tree snapshots, the capture accounting —
lives in the stream and only there.
Once a run has started (its run_id is known — supplied at launch, per §2),
an adapter can query, steer, and wait for it while it is still live. Every
command resolves the target purely by run_id through the per-user registry —
never by PID. This is also the whole supervision story for a run launched with
--detach (§2): a detached run is an ordinary run in the registry, and these
five commands are how an adapter that is no longer its parent steers it —
alongside events (§3), which reads that run's stream back without contacting it
at all:
processkit-cli inspect --run-id build-42 --json
processkit-cli cancel --run-id build-42
processkit-cli kill --run-id build-42
processkit-cli attest --run-id build-42 --json
processkit-cli wait --run-id build-42 --timeout 10mThe first four reach the live runner over the local control plane described
normatively in docs/control-plane.md; wait does not
contact the runner at all and is described in docs/registry.md,
"Waiting — wait".
inspectis read-only: it prints a snapshot (mechanism,root_pid,started_at, the currentmembers) to stdout and changes nothing — as JSON with--json(shown above), or a human-readable rendering by default.cancelends the run through the same soft-stop → grace → hard-kill teardown a--timeoutor a localCtrl-Cdrives, exiting the run withCONTROL_CANCELLED(108).killhard-kills the whole tree immediately — no soft stop, no grace — exiting the run withCONTROL_KILLED(109).attestanswers one question about the calling process: is it inside this run's container? The runner decides it from the kernel's own record of who opened the control connection — there is no--pidand no way to ask about any other process — so amemberanswer is a containment fact rather than a string the caller carried. This is what turns an adapter's "the caller belongs to run X" convention into a runner-checked invariant:verdictmemberexits0,not_a_memberexitsNOT_A_MEMBER(115) — a decided answer, deliberately not aCONTROL(103) — andpeer_identity_unsupportedexits103, the fail-closed refusal a platform that cannot name its callers gives instead of an unproven "ok" (pinattest:peer-identityat preflight, §1, to rule that out up front). The attestation is printed on stdout for every verdict, including the failing ones, and carries its ownattestation_version(§1) — this answer's contract axis, independent ofinspect'ssnapshot_version. The client reads it strictly: a reply declaring any number other than the single one this build implements is refused withCONTROL(103) andkind: "incompatible_contract"(§6) instead of being read as a verdict its sender never promised, because unlike a snapshot this answer is one an adapter gates access on — there is deliberately no read-down range. Readmechanismif you need to know how strong the containment behind amemberanswer is; nested runs, the per-mechanism scope, and why this axis has no read-down range are covered indocs/control-plane.md, "attest" and "Attestation version".waitblocks until the run is no longer live and exits0. It is the answer for an adapter that is not the runner's parent — one that restarted, or that supervises runs another process launched — and so has no child process to wait on. It prints nothing (the exit code is the answer), never touches the run, and needs no control endpoint, so it also works for a run whose transport never came up. Adding--report-outcomemakes the single-run form print one JSON object, while the--allform prints one JSON array in stable snapshot order, naming how each run ended —statusreportedwith the terminal event'scode/source/child_code, orstatusunknownwith all threenullwhen the outcome could not be established — without changing any of the exit codes below. Seedocs/registry.md, "Waiting —wait".
Each of these outputs has a published JSON Schema and golden fixture under
fixtures/schema/cli/: inspect.schema.json (the single snapshot and the
--all array), control-ack.schema.json (the cancel/kill ack and the
--all report array), attest.schema.json (the attestation), and
wait.schema.json (--report-outcome in either single-run or aggregate form).
Both mutating verbs' outcomes are also written to the target run's own
--jsonl stream (a cancelled/killed event with source
control_cancel/control_kill, and the matching terminal runner_exit), so
an adapter watching that stream sees the command take effect even without
reading the cancel/kill client's own ack.
Tearing down everything at once: --all (T-217). cancel --all / kill --all (mutually exclusive with --run-id, one of the two required) are the
aggregate counterpart to the by-run_id form above: instead of one named run
they act on every run confirmed live in a snapshot taken when the invocation
starts, applying the identical per-run mutation to each, and print a single
JSON array on stdout — one {"run_id":...,"accepted":...} entry per snapshot
target — instead of one ack. An adapter driving a full environment teardown
(e.g. before shutting down its own process) typically issues cancel --all
in place of a loop over individually-known run_ids, then wait --all /
list --json / prune to confirm the fleet is actually gone. See
docs/control-plane.md, "cancel --all / kill --all",
for the exit-code and report contract — it differs from the by-run_id form's
103 in the note right below.
Waiting for a run an adapter did not launch. The typical shape — cancel a run, then confirm it is really gone before releasing the resources it held:
processkit-cli cancel --run-id build-42 # 0: the runner acked
processkit-cli wait --run-id build-42 --timeout 30s
case $? in
0) ;; # the run is over; its own exit/JSONL say how it ended
112) ;; # still live at the deadline — the run was NOT touched
103) ;; # ambiguous run id: more than one live run uses it (§6)
esac0means "not running". It is also what an unknownrun_idreturns, on purpose: a clean exit deletes its own registry entry, so "never registered" and "already finished and cleaned up" are the same observation, and failing on the second would turn the ordinary "it finished while I was starting up" race into an error. The flip side an adapter must respect: a typo'drun_idalso returns0, so never readwait's0as proof the run existed — establish that from the launch itself or fromlist(§5).WAIT_TIMEOUT(112) is the waiter's deadline, not the run's: the run was left running and untouched, and is still going. Do not confuse it with the run's ownTIMEOUT(106) in §3's table, which means the runner tore the tree down. Retrying the samewaitis a reasonable response to a112.CONTROL(103) here means only one thing — an ambiguousrun_id(§6);waithas no runner to fail to reach.- Without
--timeout,waitblocks indefinitely. Prefer an explicit deadline in an adapter, so a supervisor never inherits an unbounded wait.
CONTROL (103) is the one exit code all five of these clients' by-run_id
form can return, and for all five the usual reason is the same: the command could
not be resolved to the single target run. Two further reasons belong to the
read-only verbs, where the target was resolved and reached and did answer. The
first is shared by both of them: a reply declaring a contract version this client
does not read is refused rather than acted on — inspect's snapshot_version (see
docs/control-plane.md, "Snapshot version: a newer runner's reply is refused, an
older one is read") or attest's attestation_version (ibid., "Attestation
version"). The second is attest's alone: it was answered
peer_identity_unsupported — the runner could not name the caller, so it declined
to decide. attest's decided negative is not a 103 at all but NOT_A_MEMBER
(115). See §6 for the concrete situations that produce a 103, those included.
cancel --all / kill --all
reuse the same code for a different reason — one or more snapshot targets failed,
not "no single target run" — see the --all paragraph above and
docs/control-plane.md.
list and prune scan the registry directly rather than reaching a specific
live run, and are the tools for an adapter that manages many runs or wants to
clean up after abrupt failures — see the normative "Discovery" and "Reaping"
sections of docs/registry.md.
processkit-cli list --json # every registered run, whatever its health
processkit-cli prune --json # reap only the confirmed-stale entrieslist --jsonprints one JSON object per registry entry (run_id, health,started_at,hint,argv_sha256,endpoint), sorted deterministically.argv_sha256andhintare the same redaction-safe command identification therun_startedevent carries (§3) — the full 64-character digest here, so an adapter can join a registry entry to the events of the run that wrote it, or group several live entries by "same command" without ever handling a command line. Both arenullon a record written before those fields existed, andhintisnullfor the common case of a command matching no known worker shape. Health islive,stale(confirmed dead — no live holder found), orunprobed(the liveness lock could not even be opened, e.g. permission denied — a distinct, additive value: liveness is unknown, never printed as the confirmed-deadstale). All three are listed, never hidden — a stale entry (a leftover from a runner that died abruptly) is exactly what an operator or adapter wants visible here, and an unprobed one is exactly the case where guessing would mislead.prune --jsondeletes only entries it can confirm are stale, printing a tally:{"pruned":N,"live":N,"unprobed":N,"orphaned_locks":N}. A live run is never touched, and an entry whose liveness could not even be probed is left in place rather than guessed at — see "The reaping safety invariant" indocs/registry.md. On unix each reaped entry also takes with it the private control-socket directory that record published, so an abruptly-killed run leaves nopkc-…litter in the temp directory either; the tally fields are unchanged (that socket is counted by its own entry'spruned). Worth scheduling if your adapter starts many runs — see "Reaping the control socket" indocs/registry.md.
Both are read-only with respect to any live run's control transport; neither carries the "could not reach the target run" failure modes of §4.
Their machine-readable shapes are published too: fixtures/schema/cli/list.schema.json
(one entry object per line) and fixtures/schema/cli/prune.schema.json (the plain
tally and, as #/$defs/dryRunReport, the --dry-run form with its candidates
list), each with a golden *.jsonl fixture beside it.
Every distinction in this section is available as a machine-readable value, not
only as prose: run any of these commands with the global --error-format json and
the failure prints one bounded JSON object on stderr whose kind names exactly the
case below (stale, unprobed, ambiguous_run_id, incompatible_contract, …).
The kind column is noted per bullet; §7 has the full contract.
- Stale registry entry. (
kind: "stale") The runner behind arun_iddied abruptly (crash,SIGKILL, a parent's Job Object terminate); its record is left behind but its liveness lock is released.inspect/cancel/kill/attestdetect this before connecting and report it as aCONTROL(103) failure with an explanatory message on stderr — never a hang, and never silently treated as live.liststill shows the entry (markedstale);pruneis what removes it. An ordinary UnixSIGTERM/SIGHUP, or a WindowsCtrl-Break/console close/logoff/system shutdown, is not in this class: the runner catches those signals/events and runs the full cancel teardown (acancelledevent, the cleanup pair,runner_exitcancelled/107, and removal of the registry entry), so stopping a run withkill <pid>(Unix) or a closed console (Windows) leaves neither a stale entry nor a surviving descendant. - Unprobeable registry entry. (
kind: "unprobed") The entry's liveness lock could not be probed at all (permission denied, a rejected symlink/reparse point, a non-regular file in its place), so nothing about the run is confirmed either way. This is the sameCONTROL(103) refusal —inspect/cancel/kill/attestact only on a confirmed-live entry — but it is reported honestly asunprobed, not as a gone runner;listshows the same entry asunprobedandpruneleaves it in place. Investigate the registry directory rather than deleting the record by hand (seedocs/troubleshooting.md). - Died mid-conversation. (
kind: "control_unreachable", or"ipc_deadline"when a bounded window elapsed instead) The registry entry read as live, but the runner exited between the liveness check and the reply reaching the client — the connect fails, or the connection closes before a complete response. Also a boundedCONTROL(103) failure, never a wedge: every wait in the control plane (connecting, and the request/response exchange) is deadline-bounded. - Ambiguous
run_id. (kind: "ambiguous_run_id") The registry does not enforcerun_iduniqueness; if more than one live entry matches, every by-run-idcommand — the read-onlyinspect,attest, andwaitincluded — fails closed withCONTROL(103) rather than guessing which entry the scan happened to return first. Keeprun_ids unique among an adapter's own concurrently-live runs (§2) to avoid this entirely. - An unreadable contract version (
inspectandattest). (kind: "incompatible_contract") The runner was reached and answered, but the reply declared a contract version this client does not read, so it was refused instead of being acted on under semantics its sender never promised. Both read-only verbs can hit this, each on its own version axis: aninspectreply carrying a control-planesnapshot_versionoutside the range this client reads — newer than the version it implements, or older than the version it still decodes — and anattestreply carrying anattestation_versionother than the single one this client reads (that axis is read strictly, with no range, since a misread membership verdict is a security answer rather than a diagnostic; §4). Also aCONTROL(103), with a message naming the version that arrived and the version — or range — this build reads. Unlike the four above it says nothing about the run's liveness: the target is registered, live, reachable, and healthy, and the run stays fully controllable, sincecancel/killacks carry no version andwait/listask the runner for nothing. Do not treat it as a lost runner or retry it; re-run the same command with a build that implements the runner's version of that contract — forinspect, its snapshot version (for a newer runner, a build at least as new as the binary that started the run); forattest, its attestation version. Seedocs/control-plane.md, "Snapshot version: a newer runner's reply is refused, an older one is read" and "Attestation version", anddocs/compatibility.md, "Machine-output schemas". - A caller the runner cannot name (
attestonly). (kind: "peer_identity_unsupported") The runner was reached and refused to decide membership, because its transport could not supply a kernel-authenticated identity for the connecting process. Also aCONTROL(103), and for the same reason as the bullet above: an answer that cannot be trusted is withheld rather than guessed — here in the safe direction, since the alternative would be reporting an unprovenmember. It says nothing about the run's liveness or about membership. Rule it out at preflight by requiringattest:peer-identity(§1); meeting it at runtime means that check was skipped or the runner is a different build. - A decided non-membership is not in this class. (
kind: "not_a_member", exitNOT_A_MEMBER115) Whenattestreports that the caller is not in the run's container, nothing failed: the target was resolved, reached, and answered. It has a code of its own precisely so an adapter can tell "the runner says no" (deny the request) from "no runner said anything" (investigate, or retry). Never fold it into the103handling above. CONTROL-class exit codes are not run outcomes. A103from the by-run_idform ofinspect/cancel/kill/attest/waitdescribes a failure on the client's side of the exchange — it could not resolve or reach a single target, or (the two read-only cases above) could not read or obtain the answer it asked for — and says nothing about how the target run itself ended (or is still running). Do not conflate it with the run-outcome codes in §3's table (106–109, or the child's own code); those come only from the run's own process exit and itsrunner_exitevent. The same separation applies toWAIT_TIMEOUT(112): it is the waiting client giving up, never the run being stopped (§4).cancel --all/kill --all's own103is the one exception where the code can coincide with some targets having genuinely been acted on — see the--allparagraph in §4.- A
--detachexit code is not a run outcome either.run --detach's0means "the run started", not "the child succeeded", and its non-zero codes mean "the run never started" — carrying the same reserved code the failure would have produced in the foreground. An adapter that branches on a detached launch's exit code as if it were the child's result will read every long-running failure as a success; the child's outcome is in the terminalrunner_exitevent (§3), reached afterwait(§4). Seedocs/exit-codes.md, "Detached runs". SETUP(111) vs.INTERNAL(104). (kind: "setup"versus"internal"; an unreadable registry narrows further to"registry") Arunthat could not write its--jsonl/--capture-dir, or open a--stdin-file, fails closed withSETUP(111) — an ordinary, usually-actionable environment problem (bad path, permissions), not a runner bug.INTERNAL(104) is reserved for a genuine invariant violation in the runner's own logic. See "Setup failures vs internal faults" indocs/exit-codes.md.
Everything in §6 is a real distinction the CLI already makes — but by default an
adapter can only read it as English on stderr, because the exit code is coarse: one
CONTROL (103) covers six of those bullets at once — eight kind values in all,
since the "died mid-conversation" bullet is two of them (control_unreachable and
ipc_deadline) and not_found, a run id the registry names nowhere, gets no bullet
of its own above. The global, opt-in
--error-format json publishes the distinction instead:
processkit-cli --error-format json inspect --run-id build-42
# stderr, exactly one line:
# {"error_version":1,"code":103,"kind":"stale","operation":"inspect",
# "run_id":"build-42","retryable":false,"message":"cannot inspect run `build-42`: …"}- Opt in wherever it is convenient. The flag is global: it parses before or
after the subcommand, and every subcommand honors it. Pin it in the preflight
like any other flag —
--require-surface inspect:--error-format(§1). - Branch on
kind(andcode), never onmessage.error_version,code,kind,operation,run_id, andretryableare the contract;messageis free text that may be reworded in any release. kindmaps onto §6.stale,unprobed,ambiguous_run_id,control_unreachable,ipc_deadline,not_found, and the two refusalsincompatible_contractandpeer_identity_unsupported— those eight are the ones that exist to split the singleCONTROL(103) — plusnot_a_member(the decided verdict,115),host_unqualified(the other decided verdict,116, and the one about the host rather than a run — §1),registry/setup/internal,wait_timeout,events_invalid,probe_incompatible, and — for a failingrun— the terminalrunner_exitevent's ownsourcespellings (spawn_error,container_error,timeout,cancelled,control_cancel,control_kill,output_overflow). Unrecognized value? Fall back tocode; the vocabulary grows additively.- stdout is untouched. The envelope is on stderr, so an adapter can leave the
flag on for every invocation without any risk to the JSON it parses from stdout —
including for a command that prints a report and then fails, such as
probe --jsonexiting 110 orinspect --all --jsonexiting 103. - The default is unchanged. Without the flag, stderr is byte-for-byte the prose every earlier release printed.
- One documented gap. clap's parse-time usage errors (exit
USAGE, 100 — an unknown flag, a malformed duration, a missing subcommand) stay human-readable in v1: they happen before the binary knows what it was asked to do, so there is no operation to name. Use the §1 preflight to establish that a flag exists before using it. Every post-parse failure is covered.
The shape is published like every other machine-readable output in this guide:
fixtures/schema/cli/error.schema.json with a golden error.jsonl beside it. The
normative field-by-field contract, the full kind table, and the retryable rule
are in docs/exit-codes.md.
An adapter reading this guide may reasonably ask for one more thing: a small file the
runner drops at the end of a run — a run --outcome-json <path>, atomically replaced
at terminal completion with the outcome, a cleanup confirmation, and artifact
locators — so the common path costs one tiny read instead of a walk over the lifecycle
stream. It was proposed, evaluated against a real adapter, and declined. This
section records why, so the question does not have to be reopened from scratch; the
durable record, including the one alternative that was deferred rather than rejected,
is ADR 0007.
- It would not remove the read it exists to remove. A supervising adapter does not
read the stream only for the terminal event. The one this decision was measured
against consumes four event types after every call —
run_started(forroot_pid,mechanism,abrupt_cleanup),members_snapshot,output_captured, andrunner_exit— and treats a missingrun_startedoroutput_capturedas a failure reason of its own. A terminal receipt carries terminal facts; the start-time facts exist only in the stream (§3). Such an adapter would gain a second artifact and keep the loop. - The read is seven lines. A foreground run emits seven events — eight with
--capture-dir, and one more again when a requested resource cap (--max-memory/--max-processes/--cpu-quota) adds its post-runlimit_evidence. Six of the seven were the whole stream untilresource_summaryjoined them unconditionally; that one line is the sole growth a caller did not opt into, and it is one line, once, at the end. Every other way the stream grows is a caller asking for more events on purpose —--capture-dir, a cap, and--snapshot-interval's extramembers_snapshotsamples.docs/schema.md, "Ordering", is the full rule. - A cheaper answer already exists, and it is the recipe in §3: the
--error-format jsonenvelope resolves exactly this ambiguity with no path to allocate, no artifact to clean up, and no second durable shape to version. - A receipt could not report its own failure. It would be written after the
child's exit code is already decided. If that write, or its atomic replace, failed,
the runner's options would be to fail the run — rewriting the child's exit code,
which
docs/exit-codes.mdforbids outright — or to say nothing, which would make the receipt's absence mean either "the runner died abruptly" or "the receipt could not be written". The property that made the idea attractive is the first thing its own failure mode would take away. The lifecycle stream is not exposed this way: it carries the whole run, so a write failure truncates it visibly rather than turning one expected artifact into a silent nothing.
What this does not claim: that every terminal read is already as cheap as it could
be. There is no first-party bounded terminal read over an arbitrary stream file today.
events --json (§3) is a whole-stream pass-through, and wait --report-outcome (§4)
reports an outcome only for a run that invocation observed live — a finished
foreground run has already deleted its own registry record, so it answers
status: "unknown" for one. If that ever becomes a measured cost rather than an
anticipated one, ADR 0007 records why the answer would be a read-side flag over
--file, reusing the shape wait --report-outcome already publishes, rather than a
second write path out of the runner.
docs/agent-workflows.md— a policy and execution strategy for automation agents that launch external tools through the runner.docs/schema.md— the normative JSONL event schema (every field, every event, versioning rules).fixtures/schema/cli/README.md— the JSON Schema documents and golden fixtures for every machine-readable output in this guide (probe,list,inspect, thecancel/killacks,prune,wait --report-outcome,attest,doctor, and the--error-format jsonfailure envelope), and the versioning decision behind them (probe,inspect,attest,doctor, and the envelope carry their own version field; the other four deliberately carry none).docs/compatibility.md— the compatibility surfaces, the pinning procedure, and the upgrade/downgrade checklists.docs/exit-codes.md— the normative reserved exit-code band and the child-fidelity rule.docs/control-plane.md— the normative local transport, wire protocol, andinspect/cancel/kill/attestbehavior.docs/registry.md— the normative registry location, record format, and staleness/reaping rules.docs/architecture.md— the map of this repository's own modules, for a contributor rather than a consumer.docs/troubleshooting.md— symptom-to-cause diagnosis for an operator, organized by what you observe rather than by call sequence.