docs: classify the frozen v1 baseline and preregister the v2 gate manifest - #79
Merged
Conversation
…ifest ## Summary - Add `review-suite/evals/v2/FAILURE-TAXONOMY.md`, classifying every material outcome across the three frozen v1 scored reports by the smallest evidenced cause, independently re-verified against each report's `per_case` data. - Add `review-suite/evals/v2/DECISION-RECORD.md`, giving #51-#57 each an explicit disposition (retain, narrow, defer) with baseline evidence, the smallest proposed schema delta, fixtures, scored corpus slice, and owner-confirmation-flagged thresholds. - Add `review-suite/evals/v2/gate-manifest.json`, preregistering the v2 corpus/grader versions, runtime/model stratum, run count/timeout/cost, deterministic invariants, proposed quality/stability thresholds, mechanism ablations, and threshold-change rules before any scored v2 output exists. - Point `review-suite/evals/README.md` at the new `v2/` directory. ## Why - #59 requires this taxonomy, decision record, and gate manifest as one reviewable candidate before any live issue body or dependency edge changes. - Every threshold is marked as a proposal requiring owner confirmation, per #59's own instruction not to assert a gate as settled from a 5-attempt sample. - This candidate changes no review prompt, schema, rubric, orchestration, or caller runtime behavior; it only classifies evidence and proposes what #51-#57 should implement.
## Summary - `DECISION-RECORD.md`'s #56 entry now states plainly that the connector-source authorization was communicated as a session-level owner instruction during this epic's delegated handoff, not (yet) as a citable tracker comment, unlike every other owner adjudication cited elsewhere in this record. - Add an explicit reconciliation with `LIMITATIONS.md` item 1's recorded exclusion of a different private connector repository from #58's frozen v1 corpus: that exclusion is scoped to the v1 corpus specifically and is unchanged; this is a separate, later-granted permission scoped to #56's own v2-cycle workflow. ## Why - review-code-change's initial pass on this candidate returned `changes_required`: the authorization claim was asserted as settled fact with no citable trail, in unacknowledged tension with the v1 baseline's own recorded exclusion of the same class of material.
…onciliation ## Summary - The #56 reconciliation bullet no longer calls the item-1-excluded repository "different" while separately claiming to leave the same-or-different question open in either direction; it now takes no position on that question at all, matching its own stated reason for staying silent. - The provenance bullet now cites only LIMITATIONS.md item 33 for "cites a specific comment URL" — item 23 makes no such claim and carries no URL. ## Why - The second review-code-change pass on this candidate returned `changes_required`: the previous fix introduced a self-contradiction (one sentence asserted the sources are different; the next sentence claimed that question was left open) and an overstated citation.
## Summary - DECISION-RECORD.md's #57 entry no longer states the stale #58 cross-reference "is corrected" in #57's body; it now says the correction is pending, to be made when this ticket's later live issue-body edit step runs, not by this documentation-only repository candidate. - The "Net effect on the dependency graph" section no longer claims #51's `blockedBy #59` edge is already removed; it now describes the actual planned behavior (leave the edge in place, matching #58's own still-present edge to closed #50) and marks the predicted net effect as something the later graph audits verify, not an already-confirmed result. - README.md (v2/ and evals/) no longer describes `audits/` as delivered content; it is now explicitly marked as not yet produced, and the dangling "native-edge mutations below" and `audits/dependency-sequencing-audit.md` references are removed or reworded to point at real, current evidence only. ## Why - The third review-code-change pass on this candidate found two related overclaims: prose asserting work as already-done (a ticket-body correction, an audit read-back, edge removal) that in fact has not happened yet and is explicitly out of scope for this documentation-only candidate. Both are the same failure class the first review round already caught once in the #56 section — settled-fact framing for something not yet verifiable.
shaug
added a commit
that referenced
this pull request
Jul 28, 2026
## Summary - DECISION-RECORD.md's #57 entry no longer says the "approved by #58" cross- reference correction is "still pending... not by this repository candidate." That fix, plus two further #58-instead-of-#59 corrections in the same body (thresholds and gate-manifest references), were already applied to the live tracker after PR #79 merged. The entry now says so, and notes scope-completeness-audit.md independently confirms no stale cross-reference remains in #57's live body. ## Why - The third review-code-change pass on this candidate found the document contradicted itself: one section called a fix "still pending," while this candidate's own audit-results section and the live tracker showed it had already happened before these audits ran.
shaug
added a commit
that referenced
this pull request
Jul 28, 2026
…aph (#80) * docs: run the three required graph audits against the live #51-#57 graph ## Summary - Add review-suite/evals/v2/audits/scope-completeness-audit.md, dependency-sequencing-audit.md, and shovel-readiness-audit.md, each an independent pass over the live GitHub issue graph after #59's approved #51-#57 body edits and native-edge read-back. - Fix two stale cross-references the scope/completeness pass found in the live tracker: #56 referenced a #52 changed-surface concept that #52's narrowing dropped, and #54 referenced a #52 risk profile that no longer exists anywhere in the graph. Neither ticket's disposition changed. - Record a summary of all three audits' findings in DECISION-RECORD.md, and point both v2/README.md and the top-level evals/README.md at the finished audit files. ## Why - #59 requires all three audits to run independently against the live graph, with findings and fixes recorded separately per pass, before the ticket can close. All three passed clean, with one flagged-but-non-blocking open question (whether #56's blockedBy #52 edge still has a real technical justification after #52's narrowing) reported for the owner rather than unilaterally resolved. * fix: sweep issue titles too, and correct a precedent timing claim ## Summary - The scope-completeness audit's cross-reference sweep checked only issue bodies. A follow-up title sweep found #52's and #53's live titles still named exactly the concepts ("risk evidence", "specialist routing") their own bodies say were dropped. Both titles are corrected via updateIssue (title only, no body change), and the audit file now records this second sweep and its findings. - The dependency-sequencing audit's #58/#50 precedent argument previously claimed #50 "closed since before #58 started," which live createdAt/ closedAt data contradicts (#58 was created before #50 closed). Replaced with the accurate timeline; the precedent's substance is unaffected. - DECISION-RECORD.md's audit summary now mentions the two title fixes alongside the two body fixes already recorded. ## Why - review-code-change's pass on this candidate returned `changes_required`: the scope-completeness audit's "no other stale cross-reference" claim was incomplete (titles unchecked), and a supporting factual claim in the dependency-sequencing audit was independently disprovable from GitHub's own timestamps. * fix: correct the remaining stale #52/#53 title references ## Summary - shovel-readiness-audit.md's #52 and #53 section headings still used the pre-fix issue titles ("Evaluate compact review coverage, impact, and risk evidence", "...and specialist routing"). Updated to match the corrected live titles. - DECISION-RECORD.md's #52 and #53 section headings had the same staleness (they predate the title fix commit); updated to match as well. ## Why - The second review-code-change pass on this candidate found the title fix from the prior commit had not propagated to two more headings inside this same delivered set, reproducing the exact staleness class the audits exist to catch. * fix: stop describing #57's already-applied prose fix as pending ## Summary - DECISION-RECORD.md's #57 entry no longer says the "approved by #58" cross- reference correction is "still pending... not by this repository candidate." That fix, plus two further #58-instead-of-#59 corrections in the same body (thresholds and gate-manifest references), were already applied to the live tracker after PR #79 merged. The entry now says so, and notes scope-completeness-audit.md independently confirms no stale cross-reference remains in #57's live body. ## Why - The third review-code-change pass on this candidate found the document contradicted itself: one section called a fix "still pending," while this candidate's own audit-results section and the live tracker showed it had already happened before these audits ran. * fix: add the updatedAt data scope-completeness-audit.md already cited ## Summary - dependency-sequencing-audit.md's live graph read-back table now includes the updatedAt column its own intro line already claimed to have queried, with the actual re-fetched timestamps for #51-#59, so scope-completeness-audit.md's pointer to "the exact updatedAt timestamps" in this table resolves to real data instead of a table that never had that column. ## Why - The fourth review-code-change pass on this candidate found a citation in scope-completeness-audit.md that pointed at data dependency-sequencing- audit.md did not actually contain, even though the underlying claim (the fixes landed when stated) was independently verifiable and true.
13 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
review-suite/evals/v2/FAILURE-TAXONOMY.md, classifying every materialoutcome across the three frozen v1 scored reports (
s1-correctness-orchestrator,s2-solution-simplicity-lens,s3-code-simplicity-lens) by the smallestevidenced cause, independently re-verified against each report's
per_casedata and
LIMITATIONS.md.review-suite/evals/v2/DECISION-RECORD.md, giving Make clean verdicts require passing validation and current-head lens evidence #51-Run the preregistered review v2 ablation and integration closeout #57 each anexplicit disposition (retain, narrow, defer) grounded in that evidence, with
the smallest proposed schema delta, fixtures, scored corpus slice, and
owner-confirmation-flagged thresholds for Add consumer/impact-traversal evidence to the shared review contract #52 and Add correctness traversal and verification-sufficiency passes #53.
review-suite/evals/v2/gate-manifest.json, preregistering the v2corpus/grader versions, runtime/model stratum, run count/timeout/cost,
deterministic invariants, proposed quality/stability thresholds, mechanism
ablations, and threshold-change rules before any scored v2 output exists.
review-suite/evals/README.mdand addreview-suite/evals/v2/README.mddocumenting the new directory.
Why
reviewable candidate before any live issue body or dependency edge changes.
caller runtime behavior; it only classifies evidence and proposes what
Make clean verdicts require passing validation and current-head lens evidence #51-Run the preregistered review v2 ablation and integration closeout #57 should implement in later tickets.
Scope note
This PR does not touch
review-suite/evals/baseline/v1/(frozen, mergedby #58) and does not edit any live issue body or native dependency edge —
those are a separate, later step per #59's own validation/delivery boundary,
applied only after this candidate is merged.
Test plan
just format— all checks passed, no reformatting neededjust lint— all checks passed (skills-ref validation of all 8 skills,plugin packaging validation)
just test— every skill'sscripts/testsplusreview-suite/scripts/tests(214 tests), all OK
review-code-changesuite (4rounds; 3 found genuine issues, all fixed; final round clean)
Refs #59