You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Make publication readiness the project's primary focus for the next seven days
(2026-07-14 through 2026-07-21). The target is an evidence-complete, frozen sota-v2 release candidate—or a precise external blocker with every safe
in-repo step finished—not more simulator scope.
The authoritative source of truth is the living checklist in docs/PUBLISH_READINESS.md.
Update it whenever evidence, a frozen decision, cost, provider state, or a
blocker changes. This issue is the compact execution tracker; #60 remains the
broader roadmap.
361 local tests, Ruff, contract canaries, calibration, archived-v1
verification, web lint/build, and empty-v2-leaderboard generation pass on the
current publication branch.
The sweep is frozen at 3 models × 4 common caps (256 / 1,024 / 4,096 /
16,384) × 3 repeats.
The primary endpoint, deterministic cap rule, retry/exclusion/stopping
policy, and minimum eight-model headline gate are committed.
Provider-dependent “uncapped” was removed before official results.
JSON mode is standardized on, reasoning is standardized off, and exact
endpoints are pinned.
All three sweep models completed a clean standardized 1,024-token smoke.
All 11 intended headline routes accept the common privacy, parameter, JSON,
reasoning-off, and bounded-output policy.
The sweep estimate is recorded: $27 planning, $32.40 with cost
contingency, $94.51 token-ceiling contingency; 8.96 projected serial API
hours, 13.44 with runtime contingency.
No official sweep or headline result exists yet; the public v2 ranking is
correctly empty.
Treat per-model p-values as descriptive unless the final family-wise
analysis justifies stronger claims.
Run power, score-weight sensitivity, leave-one-seed-out, and
extreme-episode checks.
Show output-budget curves before the ranking and show Oracle → pick-trader → best eligible model → random headroom.
Keep API and coding-harness results in separate tables.
Finish the blog from generated evidence, including the withdrawn-v1 story,
compute confound, limitations, and exact scope.
Update the README/site with the primary result, what the benchmark does and
does not measure, architecture/evaluation flow, quickstarts, cost guidance,
and links to evidence.
Add CITATION.cff, changelog/release notes, checksums, reproducibility
manifest, raw assets, and a tagged v2 release.
Check mobile, accessibility, commands, internal links, and clean-clone
behavior.
Obtain an independent final read and, if possible, one outside
clean-clone reproduction.
Decide before writing claims whether a salted private-panel precommit/run
is feasible; otherwise scope v2 explicitly to reproducible public-panel
evidence.
Release gate
Do not publish a headline ranking unless all are true:
Frozen benchmark and compute policies are fingerprinted and reproducible.
Every headline row is strictly eligible and compute-comparable.
At least eight pre-registered models pass.
Raw evidence is available and hash-linked from compact artifacts.
Failures, exclusions, uncertainty, compute, cost, and limitations are
visible.
Claims stay inside this synthetic model-plus-scaffold condition.
CI, artifact validation, site build, commands, links, and clean-clone
checks pass.
No v3 behavior has leaked into the frozen v2 lane.
Daily operating rhythm
Start each work session with this issue and the living document.
End each session by checking completed items and linking durable evidence.
Record provider/quota drift immediately instead of burning repeat calls.
Keep the next unblocked critical-path item explicit.
Update cost/runtime estimates whenever observed data materially changes
them.
Publishing a paper; the deliverables are the benchmark, GitHub release/site,
reproducible artifacts, and a rigorous blog post.
Closes when v2 is published, or when the only remaining work is an explicitly
documented external run/review blocker and every safe in-repo task is complete.
Goal
Make publication readiness the project's primary focus for the next seven days
(2026-07-14 through 2026-07-21). The target is an evidence-complete, frozen
sota-v2release candidate—or a precise external blocker with every safein-repo step finished—not more simulator scope.
The authoritative source of truth is the living checklist in
docs/PUBLISH_READINESS.md.Update it whenever evidence, a frozen decision, cost, provider state, or a
blocker changes. This issue is the compact execution tracker; #60 remains the
broader roadmap.
Focus rule for the week
sota-v3work parked.official model runs begin.
score.
spend ceiling.
Current state
main.mainafter independent review and full finding disposition.verification, web lint/build, and empty-v2-leaderboard generation pass on the
current publication branch.
16,384) × 3 repeats.
policy, and minimum eight-model headline gate are committed.
endpoints are pinned.
reasoning-off, and bounded-output policy.
contingency, $94.51 token-ceiling contingency; 8.96 projected serial API
hours, 13.44 with runtime contingency.
correctly empty.
1. Land the publication pipeline
exact final SHA.
main.estimate, contract checks, and web build.
Exit:
maincontains the frozen protocol and safe paid-run path.2. Run and interpret the output-budget sweep
capabilities, and route health immediately before the run.
a conservative guard, not automatic authorization.
scripts/run_publication_matrix.py sweep --max-spend-usd <approved>.summaries, and account-cost deltas outside git.
poor scores, or surprising valid cells.
results/analysis/output-budget-sweep.json.cost, and latency for every model/cap cell.
16,384 plus compute-elastic curves if no lower cap saturates.
config/sota_v2_lane.jsonwith the decision, date, artifact hash,rationale, and limitations before any full-panel result.
Exit: one compute policy is frozen from complete sweep evidence.
3. Run and validate the headline model panel
explicit spend approval.
models after results are visible.
and provenance telemetry.
pre-registered infrastructure cases.
sota-v2.results/diagnostics/; never weakenthe headline gate.
appear.
public traces as release assets.
Exit: the generated leaderboard contains a meaningful, compute-comparable,
strictly eligible panel.
4. Analyze, write, and package
compute, cost, failures, repairs, and per-mechanic outcomes.
analysis justifies stronger claims.
extreme-episode checks.
pick-trader→ best eligible model →randomheadroom.compute confound, limitations, and exact scope.
does not measure, architecture/evaluation flow, quickstarts, cost guidance,
and links to evidence.
CITATION.cff, changelog/release notes, checksums, reproducibilitymanifest, raw assets, and a tagged v2 release.
behavior.
clean-clone reproduction.
is feasible; otherwise scope v2 explicitly to reproducible public-panel
evidence.
Release gate
Do not publish a headline ranking unless all are true:
visible.
checks pass.
Daily operating rhythm
them.
Out of scope this week
reproducible artifacts, and a rigorous blog post.
Closes when v2 is published, or when the only remaining work is an explicitly
documented external run/review blocker and every safe in-repo task is complete.
Related: #60, #58, #59, #61, #62, #46.