These files are public calibration fixtures for the v2 workflow-resilience scorer. They are not contender results and do not prove that any model or workflow is superior.
workflow-resilience-pilot.scenario.json— pilot scenario that combines a v1 artifact score with scorer feedback, context-wipe recovery, probe-only, stale, impossible, and final stop-decision phases.workflow-resilience-smoke.scenario.json— smaller smoke fixture for the first scorer implementation.
pilot-good-run-record.json— ranked fixture with complete generic outputs, replayable public refs, correct trap handling, and intentionally weak artifact quality imported from the sample v1 scorer JSON.pilot-weak-ranked-run-record.json— ranked fixture with lower stop/replay scores while still satisfying the public contract.pilot-wrong-stop-run-record.json— ranked fixture that keeps required outputs but makes the wrong final continue decision under trap phases.pilot-missing-output-run-record.json— rank-block fixture missing a required context-wipe handoff output.pilot-private-path-run-record.json— rank-block fixture with a private/local v1 scorer input ref.
paired-private-pilot.campaign.json— draft paired-campaign config for the next private pilot. It freezes Lane A as the naked model baseline, Lane B as the Open Scaffold ledger/analyze lane, and Lane C as disabled until a real controller exists. It is a protocol fixture, not a completed run result.private-pilot-result-template.md— Phase 2 recording template for per-pair, per-lane generation evidence, v1-to-v2 run-record mapping, visual package refs, context-wipe handoff, and bounded decision rules.
The canonical Stage 1 pre-pilot smoke command is:
python3 scripts/run_v2_stage1_smokes.pyIt regenerates the scorer-result fixtures to a temp directory, checks them against the committed JSON, validates the result spine and campaign fixture, and creates a deterministic visual/replay package from the fixed campaign seeds. The visual package is packaging evidence only; it is not a visual-quality judgment.
To inspect the package, write it to an empty directory:
python3 scripts/run_v2_stage1_smokes.py --visual-out <empty-output-dir>mkdir -p v2/examples/results
for name in good weak-ranked missing-output private-path wrong-stop; do
cargo run --quiet -p m2000-v2-conformance -- \
v2/examples/workflow-resilience-pilot.scenario.json \
v2/examples/pilot-${name}-run-record.json \
--json-out v2/examples/results/pilot-${name}-result.json
done
python3 scripts/render_results.py --checkA valid calibration pack should include both ranked and rank-blocked examples so future scorer changes cannot collapse v2 into score-only artifact deltas.
Validate campaign fixtures with:
python3 scripts/validate_v2_campaigns.py