fix(overnight-multi-issue-implementation): drop two figures 1.3.0 could not source (1.3.1) - #26
Merged
Merged
Conversation
…ld not source (1.3.1)
1.3.0 shipped, inside the "never truncate a findings payload" lesson, a figure
of exactly the kind the same release's reviewer line exists to catch. Found by
checking it against the artifact, which is the only way any of these are ever
found.
1. "in a 553-line pre-registration with no power statement anywhere"
The pre-registration is 1,109 lines, not 553. Both numbers were true: 553
was the document the reviewer read, and the PR then had five blocking
findings fixed on the branch before it merged, roughly doubling it. So the
figure was accurate when the finding was written and wrong about the
artifact as it now stands, with no vintage attached to tell a reader which
is which -- in a document the skill's readers cannot open.
Dropped rather than dated. The load-bearing claim is "no power statement
anywhere"; the length was decoration, and dating it would only invite the
same rot again. Now reads "in a pre-registration that had no power
statement anywhere" -- past tense, because the document has had one since
the fix.
2. "About one in four, on documents."
Retracted at the source. The host repo's own later audit reopened the count
of five: a sixth instance surfaced in a PR body, and two of the five were
documents written from scratch, where "the defect it existed to repair" is
a stretch. Its ruling is to treat this as "common enough to budget for, not
a measured rate", and quoting a rate the source has withdrawn is the same
defect one layer up.
The count of five stays, attributed as the run's own count and immediately
qualified. What was always load-bearing is untouched and now stated
plainly: the diff never catches these, only re-derivation does.
Every other figure in the two blocks was checked against the run's handoffs
and HOLDS, so nothing else moved:
six findings / five arrived the run's PR record: "a three-lens review
returned six blocking findings and the
orchestrator truncated the payload so five
arrived"
"about five times finer" +/-0.03 margin against a smallest detectable
difference of ~0.17, recorded as held on
re-derivation
server 269 -> 314 three separate handoffs record that exact
transition
slice(0, 9000) the mechanism itself, not a measurement
Also checked and left alone: eight replayed PRs, two of three board tasks
claimed, and the run scale in the worked example (17 items, 4 workflows, 91
subagents, ~11 hours, 20 PRs). Nothing contradicts them.
Versions, all four places, for a two-phrase change -- the repo's known failure
mode is a content change that never ships because no version moved:
SKILL.md frontmatter 1.3.0 -> 1.3.1
plugins/.../.claude-plugin/plugin.json 1.3.0 -> 1.3.1
.claude-plugin/marketplace.json 1.3.0 -> 1.3.1
VERSION (bundle) 1.3.0 -> 1.3.1
README carries the correction too, as its own version-history entry naming
both figures. A correction that reaches only the SKILL leaves the changelog
asserting the withdrawn rate.
No description changed: the gate reports 1,463 chars with 73 to spare and
--compare reports 8 triggers unchanged, none dropped or narrowed. SKILL.md
742 -> 747 lines. Gates: description-cap exit 0, leak exit 0,
validate_plugins exit 0, listing budget fits at 1M with 9,496 chars spare.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013JVMGELnoA2vM56Va4gLXE
wan-huiyan
force-pushed
the
fix/unsourceable-figures-1.3.1
branch
from
August 6, 2026 11:05
717f53e to
cca187b
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #25. That release shipped, inside the "never truncate a findings payload" lesson, a figure of exactly the kind the same release's reviewer line exists to catch — a number quoted from a document, with no vintage, that a reader of the skill cannot re-derive. It was found by opening the artifact, which is the only way any of these are ever found.
1. "in a 553-line pre-registration" — dropped
The pre-registration is 1,109 lines, not 553.
Both numbers were true. 553 was the document the reviewer read; the PR then had five blocking findings fixed on the branch before it merged, roughly doubling it. So the figure was accurate when the finding was written and wrong about the artifact as it stands, with nothing to tell a reader which is which.
Dropped rather than dated. The load-bearing claim is "no power statement anywhere" — the length was decoration, and dating it would only invite the same rot again in a document the skill's readers cannot open. Now reads "in a pre-registration that had no power statement anywhere", past tense, because the document has had one since the fix.
2. "About one in four, on documents." — retracted, because the source retracted it
This one the audit did not expect. The host repo's own later audit reopened its count of five: a sixth instance surfaced in a PR body, and two of the five were documents written from scratch, where "the defect it existed to repair" is a stretch. Its ruling, verbatim, is to treat this as "common enough to budget for, not as a measured rate — the useful part is not the number, it is that the diff never catches it."
Quoting a rate the source has withdrawn is the same defect one layer up, so the rate is gone. The count of five stays, attributed as the run's own count and immediately qualified. What was always load-bearing is untouched and now stated plainly: the diff never catches these; only re-derivation does.
Every other figure was checked and holds
server 269→314slice(0, 9000)Also checked and left alone, with nothing contradicting them: eight replayed PRs, two of three board tasks already claimed, and the run scale in the worked example (17 items, 4 workflows, 91 subagents, ~11 hours, 20 PRs).
One incidental note on why dating the figure would not have helped: the pre-registration measures 1,109 lines on one checkout and about 1,100 on another. A nine-line spread between two working copies of the same document, on the same day, is the argument for dropping the count rather than pinning it.
Versions — all four places, for a two-phrase change
A one-line content change still needs the bump; this repo has come back for that three times (#20, #21, #24). Bumping
VERSIONcuts a v1.3.1 release, same as #25.The README carries the correction as its own version-history entry, naming both figures. A correction that reaches only the SKILL leaves the changelog still asserting the withdrawn rate — and the changelog is a rendered surface.
Gates
Listing budget,
--context 1000000:needed 30,504 chars (~7,626 tok) · budget 40,000 (~10,000 tok) · fits, 9,496 chars to spare.No description changed —
--comparereports1,463 → 1,463 chars · 8 triggers unchanged · No trigger dropped or narrowed.SKILL.md742 → 747 lines.🤖 Generated with Claude Code
https://claude.ai/code/session_013JVMGELnoA2vM56Va4gLXE