diff --git a/PLAN.md b/PLAN.md index 58c27d7..1792c7e 100644 --- a/PLAN.md +++ b/PLAN.md @@ -40,8 +40,12 @@ every commit is world-readable and every merge to `main` is a publication act. - The calibration protocol family now has separate E4 grader-validation, E1 Lane A A/A session-position, and E2 operating-characteristic candidate SAPs. They are not preregistered, authorize no observations, and select no powered - method. Their counts, thresholds, grids, and consequences remain review - candidates until separately frozen and registered. + method. The 2026-07-17 design decisions are incorporated: E1 uses 120 + cross-arm pairs with conservative stage-localized falsification, E4 certifies + overall adjudicated-positive/negative populations under asymmetric thresholds, + and E2 uses rule-relative nuisance control with exact or simultaneous-bound + admission. Exact corpus strata, immutable implementations, schemas, and other + named freeze inputs remain registration blockers. - Contract B `2.0.0-draft.1` is an immutable, repository-only normative draft beside v1. It binds synthetic registration, reveal, schedule, attempts, attrition, copied observation provenance, analysis, decision, and report diff --git a/README.md b/README.md index 9814a77..9e4de54 100644 --- a/README.md +++ b/README.md @@ -110,6 +110,10 @@ Neither is a current statistical result. The smoke artifact retains summary valu raw per-attempt outcomes needed for independent interval recomputation. Their recorded values are historical artifacts, not evidence for model selection. +The candidate calibration family incorporates the 2026-07-17 statistical design decisions but is +still unregistered and non-authorizing. Its remaining freeze fields and operational prerequisites +must be completed outside any observation before a study can begin. + ### Verify locally Install the exact Quarto release named in `.quarto-version` (on macOS, diff --git a/STATUS.md b/STATUS.md index 879f640..d5279df 100644 --- a/STATUS.md +++ b/STATUS.md @@ -16,10 +16,10 @@ ready for a powered scientific promotion decision or production use. | Item | Evidence | Status | |---|---|---| -| Public site | Render run `29614595319` and Pages run `29614595347` passed for current published baseline `6aabf7f` | published | +| Public site | Render run `29657510137` and Pages run `29657510147` passed for current published baseline `5423caa` | published | | Core version | `pyproject.toml` and frozen core identify `0.2.0`; wheel containment is tested | repository-versioned, not distributed — the version names code in this repository, not a release anyone can install from a registry | | Experimental calculations | Committed synthetic results regenerate byte-for-byte | decision input only — the numbers inform, but do not make, the pending method decision | -| Calibration SAP family | Separate E4, E1, and E2 candidate protocols define reviewable objectives, counts, failure rules, and claim boundaries | not preregistered; no observations authorized; numerical choices remain candidates | +| Calibration SAP family | Separate E4, E1, and E2 candidate protocols incorporate the 2026-07-17 statistical decisions and retain explicit freeze blockers | not preregistered; no observations authorized; exact corpus strata, implementations, schemas, and operational bindings remain incomplete | | Contract B v2 draft | Exact schema, verifier, renderer, safety scanner, fixture, report, and delivery manifest are digest-bound (every implementation byte is hash-pinned) | synthetic conformance only — the draft evidence contract is validated against invented data, not real campaigns | | Independent refusal | The public verifier derives `NOT_EVALUABLE` from the synthetic fixture and rejects mutations — it recomputes the result from frozen evidence bytes and refuses any altered copy | verified for the fixture | | Publication scanning | Evidence inputs are recursively scanned in the parameterized report path; Contract B checks its bound public payload separately | path-scoped; no whole-repository scan — leak scanning covers the report inputs, not every file in the repository | @@ -73,5 +73,5 @@ conformance vector are unchanged. ## Provenance -- Status date: 2026-07-17 -- Baseline reviewed for this candidate update: `6aabf7f32e3d824ee63f2eef3a761f29be4e2d4c` +- Status date: 2026-07-18 +- Baseline reviewed for this candidate update: `5423caa370664f1e34101e8e812397f6a9608e1a` diff --git a/docs/preregistrations/calibration-sap.md b/docs/preregistrations/calibration-sap.md index 480f4bd..8bfad4a 100644 --- a/docs/preregistrations/calibration-sap.md +++ b/docs/preregistrations/calibration-sap.md @@ -2,43 +2,59 @@ ## Metadata -- **Date:** 2026-07-17 -- **Version:** 1.1.0-candidate.1 -- **Status:** Candidate protocol — not preregistered; no observations authorized. +- **Decision date:** 2026-07-17 +- **Last updated:** 2026-07-18 +- **Version:** 1.2.0-candidate.2 +- **Status:** Candidate protocol — design decisions incorporated; not preregistered; no observations authorized. ## Purpose -This document is the index and common claim boundary for three separate candidate statistical +This document is the index and common public claim boundary for three separate statistical analysis plans: 1. [E4 grader validation](e4-grader-validation-sap.md) 2. [E1 Lane A A/A session-position study](e1-lane-a-aa-sap.md) 3. [E2 operating-characteristics simulation](e2-operating-characteristics-sap.md) -The former umbrella draft combined three different studies and incorrectly described the E1 -objective as estimation of within-session correlation. The replacement family uses Lane A `m = 1` -by construction: one task pair per fresh session. E1 is therefore a determinism/falsification study, -not a correlation-estimation study. +Lane A uses `m = 1` by construction: one task pair per fresh session. E1 is therefore an +exact-identity falsification study, not a correlation-estimation study. E4 validates the grader +against independently adjudicated human labels. E2 evaluates candidate decision rules and later +inherits accepted E4 and E1 inputs for sensitivity analysis. -## Common rules +## Common public rules - E4, E1, and E2 require separate reviewed registrations and separate result artifacts. -- Planned values in these documents are design candidates, not produced observations. -- Every initiated E1 attempt is retained as exactly one terminal attempt row. Failure, timeout, - resource exhaustion, receipt failure, and harness failure are outcomes; none is silently dropped. -- No whole-attempt automatic retry is allowed. A separately authorized rerun must use a new attempt - ordinal and a new registration rather than replacing the original record. -- Each gateway call has a planned provider retry limit of zero. A client invocation may contain - multiple disclosed calls; all calls must be included in its closure record. -- No fallback model, direct provider path, or unregistered schedule substitution is allowed. -- A missing registration, byte binding, assignment row, owner receipt, invocation closure, terminal - attempt row, or required analysis artifact makes the affected study `NOT_EVALUABLE`. +- Values in these documents are frozen design inputs only after their remaining named fields are + completed and registered. They are not produced observations. +- Every initiated attempt is retained as exactly one terminal attempt row. Failure, timeout, + resource exhaustion, provenance failure, and harness failure are outcomes; none is silently + dropped. +- No whole-attempt automatic retry is allowed. A separately authorized rerun uses a new attempt + ordinal under a new registration rather than replacing the original record. +- No unregistered schedule, rule, corpus, or serving-profile substitution is allowed. +- A missing registration, byte binding, assignment row, required provenance record, terminal row, + or analysis artifact makes the affected study `NOT_EVALUABLE`. +- Every study binds one versioned serving-profile digest. Public analysis treats that profile as an + opaque registered input; deployment-specific values, credentials, and authority remain outside + this public methodology. - Calibration results cannot promote a system, select a powered method, or settle a minimum practically important benefit. Those require a later coupled ruling and a new protocol. +## Dependency order and phasing + +E4 and E1 may proceed independently after their own registrations. E2 has two phases: + +1. the base operating-characteristics grid uses frozen synthetic distributions and may execute + without E4 or E1 observations; and +2. grader-error and apparatus-sensitivity cells execute only after accepted E4 and E1 result + packets exist. + +The sensitivity transformation and input schema must be frozen before the E4/E1 results are +observed. This prevents later results from tuning the sensitivity analysis. + ## Current readiness -These documents are reviewable design inputs only. Before any study can be registered, reviewers -must resolve the candidate decisions called out in each SAP, freeze the machine-readable inputs and -implementations by digest, approve the evidence/report contracts, and establish operational -registration and observation authority. +The statistical decisions from the 2026-07-17 review are incorporated below, but the family is not +registered. Registration remains blocked on the incomplete E4 corpus-stratum fields, exact +machine-readable artifacts and implementations, reviewed evidence/report schemas, and separately +established operational authority. Nothing in this public family authorizes execution. diff --git a/docs/preregistrations/e1-lane-a-aa-sap.md b/docs/preregistrations/e1-lane-a-aa-sap.md index 1ea525b..d83ae41 100644 --- a/docs/preregistrations/e1-lane-a-aa-sap.md +++ b/docs/preregistrations/e1-lane-a-aa-sap.md @@ -2,107 +2,121 @@ ## Metadata and claim boundary -- **Date:** 2026-07-17 -- **Version:** 0.1.0-candidate.1 -- **Status:** Candidate protocol — not preregistered; no observations authorized. -- **Study type:** Local-lane calibration; structurally non-promotable. +- **Decision date:** 2026-07-17 +- **Last updated:** 2026-07-18 +- **Version:** 0.2.0-candidate.2 +- **Status:** Candidate protocol — design decisions incorporated; not preregistered; no observations authorized. +- **Study type:** Local-lane exact-identity falsification; structurally non-promotable. -All counts, seeds, and rules below are **planned candidate values**. They are not observations and -do not establish a powered-study replicate policy. +The candidate design counts and rules below are protocol inputs, not observations. They establish +no powered-study replicate policy. -## Objective +## Objective and estimand -E1 attempts to falsify exact outcome invariance when two blinded arms resolve to the same content -digest and task pairs are isolated by fresh sessions. Lane A uses **`m = 1`**: one task pair per -fresh session. The design removes within-session clustering from the later task-level comparison by -construction; E1 does not estimate an intra-session correlation and does not test a null hypothesis -that such a correlation is zero. +E1 attempts to falsify exact cross-arm identity when two blinded arms receive byte-identical +registered scientific invocation payloads under the same versioned serving profile. Lane A uses +**`m = 1`**: one task pair per fresh session. E1 does not estimate an intra-session correlation and +does not claim general determinism outside the registered apparatus. -E1 reports whether task success and exact outcome digests remain identical across A/A arms and -registered session positions. Latency and resource measures are diagnostics only and never turn a -divergent outcome into an invariant one. +The primary estimand is the cross-arm divergence frequency among the 120 planned session pairs. +Cross-position comparisons of the same task across rounds are a separately reported diagnostic; +they are not pooled into the primary denominator. Latency, resource use, and warmup behavior are +diagnostics and never turn a divergence into identity. -## Planned design and counts +## Design to register and counts -- 40 registered task pairs, with 10 in each of four fixed task classes. +- 40 task pairs, with 10 in each of four fixed task classes. - 3 rounds, each containing one fresh session for every task pair. - 120 sessions total: `40 task pairs x 3 rounds`. +- 120 primary cross-arm comparisons. - 2 blinded arm attempts per session, yielding 240 terminal arm attempts total. -- Each arm attempt has a separate editing-client descriptor and a separate gateway invocation with - open, zero-or-more call, and closure receipts. -- Both arms must bind the same resolved content digest, decoding configuration, task bytes, harness, - grader, prompt, identity-domain preimage, and resource policy. +- Each arm attempt has an isolated workspace, fresh client process, separate invocation record, and + separate terminal row. +- Both arms bind the same task bytes, registered scientific invocation payload bytes, harness, + grader, prompt, identity-domain preimage, resource policy, and serving-profile digest. -The number of model calls is not fixed at 240. A client invocation may make multiple gateway calls, -all of which must appear in invocation closure. Zero calls are valid only for a retained terminal -pre-dispatch failure. +The exact scientific request-body bytes are hashed immediately before transmission and must match +between arms. Transport-envelope fields that must differ, such as unique record identifiers, must +remain outside that body, be enumerated before registration, and be bound separately. No +post-transmission canonicalization may erase a request-body difference. An unregistered byte +difference makes the comparison `NOT_EVALUABLE`. -## Planned assignment schedule +## Assignment schedule -The candidate schedule seed is the UTF-8 string `lane-a-aa-schedule-v1`. Sort registered opaque task +The schedule seed is the UTF-8 string `lane-a-aa-schedule-v1`. Sort registered opaque task identifiers by `SHA-256(seed || NUL || task_id)`. Use that base order in round 1 and rotate it by 13 -and 27 positions in rounds 2 and 3. The machine-readable schedule must bind the resulting task, -round, planned-position, and arm-order rows before registration. +and 27 positions in rounds 2 and 3. The machine-readable schedule binds task, round, +planned-position, and arm-order rows before registration. Within each round, arm order is determined by the parity of `round_index + planned_position`, where -round indices are 0, 1, and 2 and positions are 1 through 40. This yields exactly 20 sessions in each -arm order per round and 60 in each arm order overall. - -The execution record must also store the **actual chronological session position**. Analysis uses -actual position and reports any deviation from the registered order; a synthetic early/middle/late -label is not a substitute. - -## Session isolation and execution - -Before each task-pair session, create isolated workspaces and fresh client processes for both arms. -Each arm gets its own invocation and workspace; neither arm may read the other's artifacts. Close -both invocations and emit both terminal attempt rows before beginning the next session. Record -machine-load, model-residency, cache, process, harness, and verification-subprocess diagnostics that -the registered observation contract can actually support. Do not claim an engine restart unless it -is observed and receipt-bound. - -The provider retry limit is zero for every gateway call. There is no automatic whole-attempt retry. -Client-generated additional calls are disclosed within the same invocation and do not create a new -scientific attempt. Every failure remains in the attempt ledger. - -## Analysis and falsification rule - -For each task and round, compare the two arms' `task_success` and exact `outcome_digest`. For each -task across its three actual positions, compare the same fingerprint. Report: - -- cross-arm and cross-position divergence counts and exact task/round locations; -- terminal status and failure-mode counts by arm order, round, class, and actual position; -- latency and resource summaries by the same strata, labeled diagnostic; and -- all schedule deviations, receipt failures, and incomplete joins. - -The candidate E1 result has four states, applied in the following precedence order: - -- `INVARIANCE_NOT_REFUTED` when all 120 sessions and 240 arm attempts have complete terminal - accounting, every required owner join and digest binding validates, no unplanned terminal - apparatus failure occurs, and there is zero cross-arm and zero cross-position divergence; -- `INVARIANCE_REFUTED` when the evidence is complete and valid but at least one fingerprint - diverges; -- `APPARATUS_NOT_ADMISSIBLE` when fingerprints do not diverge but complete, valid evidence contains - an unplanned timeout, resource, harness, or client failure; or -- `NOT_EVALUABLE` when missing or invalid registration, schedule, binding, receipt, closure, or - terminal accounting prevents the falsification analysis from being trusted. - -Thus a valid negative observation is reported as an invariance falsification or an apparatus -admission failure, not discarded as missingness. -This exact-invariance rule is a candidate requiring formal review; a clean result does not prove -invariance outside the registered apparatus and makes no population-level claim about a correlation -parameter. - -## Hard stops and consequences - -Stop dispatch on a content-digest mismatch, fallback, direct unregistered transport, broken receipt -chain, schedule mutation, evidence-integrity failure, or loss of session isolation. Already initiated -attempts remain reportable terminal rows. A validly observed timeout, resource exhaustion, harness -error, or client error is not dropped or replaced and blocks apparatus admission; absent fingerprint -divergence it yields `APPARATUS_NOT_ADMISSIBLE`. A broken evidence chain yields `NOT_EVALUABLE`. - -An evaluable E1 result is evidence only about the frozen local apparatus and schedule. It cannot -validate a remote lane, select an enforcing test, define the powered replicate unit, or authorize -promotion. -Before registration, reviewers must freeze the executable schedule, wall-time bound, zero-dollar -resource budget, stop implementation, evidence schema, and report schema by exact digest. +round indices are 0, 1, and 2 and positions are 1 through 40. This yields exactly 20 sessions in +each arm order per round and 60 in each arm order overall. + +The execution record also stores the **actual chronological session position**. Analysis uses +actual position and reports every deviation from registered order. + +## Serving-profile and recomputation requirements + +Every E1 record binds one registered serving-profile digest. That profile must structurally enforce +serialized execution, independent prompt computation, transmitted decoding parameters, adequate +context capacity, quiescence, and the registered resource bounds. Exact deployment values and +mechanisms are not part of this public document. + +A scientific `INVARIANCE_NOT_REFUTED` requires authoritative, request-bound evidence that the +second arm independently recomputed its prompt. Timing is corroborative only. If the registered +apparatus cannot provide the authoritative recomputation record, E1 is `NOT_EVALUABLE`. + +Each session begins with one standardized, non-task warmup invocation whose result is discarded and +recorded outside the 240 scientific attempts. The warmup must not provide reusable task-prefix +state. Cross-arm divergence is reported as a possible first-call/warm-state effect; cross-position +divergence is reported separately as a possible position or state-leakage effect. + +## Stage localization + +Every arm records this digest chain: + +`fixture -> invocation -> transcript -> workspace_post -> diff -> outcome` + +For every divergence, report the earliest differing stage. `outcome_digest` is the sensitive +end-check; `task_success` is the estimand-level check. A transcript-stage classification is allowed +only when fixture and registered invocation bytes match exactly and every required upstream digest +and recomputation record validates. + +The non-causal localization label is +`ENGINE_STAGE_DIVERGENCE_UNDER_REGISTERED_APPARATUS`. It identifies the earliest observed stage +under one registered apparatus; it does not prove that an engine algorithm was the sole cause. +Missing or ambiguous upstream evidence yields `NOT_EVALUABLE`, never this label. + +## Falsification rule and state precedence + +The terminal state is assigned in this conservative precedence order: + +1. `NOT_EVALUABLE` when missing or invalid registration, schedule, binding, provenance, + recomputation, or terminal accounting prevents trusted analysis; +2. `APPARATUS_NOT_ADMISSIBLE` when otherwise complete evidence contains any unplanned timeout, + resource, harness, or client failure, including a failure that co-occurs with divergence; +3. `INVARIANCE_REFUTED` when all evidence is complete, the apparatus is admissible, every upstream + scientific byte and digest matches, and at least one of the 120 primary cross-arm comparisons + cleanly diverges; the divergence carries the localization label above; or +4. `INVARIANCE_NOT_REFUTED` when all 120 sessions and 240 attempts are complete and admissible and + all 120 primary cross-arm comparisons are identical. + +Cross-position results are emitted as a separate diagnostic result and do not change the primary +cross-arm denominator or silently acquire a causal label. + +With zero clean cross-arm divergences, the nominal rule-of-three upper detection bound is +approximately `3 / 120 = 2.5%`. This does not prove invariance, cannot exclude rarer divergence, and +is further limited by repeated-task dependence because the 120 pairs arise from 40 tasks across +three rounds. + +## Hard stops and remaining freeze work + +Stop dispatch on an unregistered byte difference, serving-profile mismatch, schedule mutation, +loss of isolation, evidence-integrity failure, or missing authoritative recomputation evidence. +Already initiated attempts remain terminal rows. + +Before registration, reviewers must freeze the executable schedule, canonical scientific-payload +rule, serving-profile digest, recomputation-record schema, wall-time and resource bounds, stop +implementation, evidence schema, classifier implementation, and report schema by exact digest. +This public SAP does not establish operational authority or reveal deployment-specific apparatus +values. diff --git a/docs/preregistrations/e2-operating-characteristics-sap.md b/docs/preregistrations/e2-operating-characteristics-sap.md index ccc4a9f..a30976f 100644 --- a/docs/preregistrations/e2-operating-characteristics-sap.md +++ b/docs/preregistrations/e2-operating-characteristics-sap.md @@ -2,75 +2,103 @@ ## Metadata and claim boundary -- **Date:** 2026-07-17 -- **Version:** 0.1.0-candidate.1 -- **Status:** Candidate protocol — not preregistered; no observations authorized. +- **Decision date:** 2026-07-17 +- **Last updated:** 2026-07-18 +- **Version:** 0.2.0-candidate.2 +- **Status:** Candidate protocol — design decisions incorporated; not preregistered; no observations authorized. - **Study type:** Deterministic synthetic simulation; no model or provider execution. -Every value below is a **planned candidate value**. E2 compares candidate decision rules; it does -not select one, settle R1/R2/R4, or produce model-performance evidence. +E2 compares candidate decision rules. It does not select a powered method or produce +model-performance evidence. -## Objective +## Objective and phases -E2 estimates operating characteristics of exact, hash-bound candidate rule implementations under -the finite 40-task, four-class design. It evaluates Type I error, power, missingness, grader error, -heterogeneity, and residual apparatus sensitivity before any powered protocol can be considered. +E2 characterizes finite-sample Type I error, power, conservativeness, missingness behavior, grader +error, heterogeneity, and residual apparatus sensitivity for the finite four-class design. -No rule is currently accepted. E2 cannot run until the candidate implementations, estimands, -nuisance handling, and simulation distributions are frozen by exact digest. A result for an -unfrozen or approximate stand-in does not apply to the later powered rule. +E2 is divided into: -## Candidate scenario grid +1. a base grid using frozen synthetic paired-outcome distributions, perfect grading, and exact + apparatus identity; and +2. sensitivity cells using accepted E4 grader-error matrices and accepted E1 apparatus behavior. -The proposed grid crosses: +The base phase may execute after its own registration without E4/E1 results. The sensitivity +transformation, input schema, and boundary cases must be frozen before E4/E1 results are observed; +the sensitivity cells execute only after those accepted inputs exist. + +## Candidate rule set + +The leading candidate is the `delta_0 = 0` exact conditional paired binary +(McNemar/sign-equivalent) test. Its estimand, fixed horizon, two-sided or one-sided direction, +missingness policy, and exact implementation must be frozen by digest. + +A bounded-mean betting procedure and any positive-margin rule remain held design slots. Neither may +be implemented or simulated until it has an explicit estimand, horizon or stopping rule, +missingness policy, nuisance treatment, and immutable implementation. + +Each candidate rule has its own null/alternative partition. For example, an effect cell may be an +alternative for a zero-margin rule and a null cell for a positive-margin rule. + +## Base scenario grid + +The proposed feasibility grid crosses: - total task-pair counts `N` in `{20, 40, 60}` with equal planned class weighting; - incumbent success probabilities in `{0.30, 0.50, 0.70, 0.90}`; - mean candidate-minus-incumbent effects in `{0.00, 0.025, 0.05, 0.10, 0.20}`; - effect patterns: homogeneous, balanced opposing class effects, sparse benefit, and one-class harm; -- terminal missing/failure rates in `{0.00, 0.01, 0.05, 0.10}` under registered arm-symmetric and - arm-asymmetric mechanisms; -- grader-error matrices from the accepted E4 result plus registered boundary sensitivity cases; and -- exact apparatus invariance plus registered divergence cases derived from the accepted E1 result. +- terminal missing/failure rates in `{0.00, 0.01, 0.05, 0.10}` under arm-symmetric and + arm-asymmetric mechanisms; and +- the feasible paired-outcome discordance nuisance domain for every rule-relative null cell. + +`N = 60` is a hypothetical feasibility point. The current suite is frozen at 40; a real study with +`N > 40` requires a separately governed suite expansion under the per-class sampling frame. Impossible probability combinations are omitted by an explicit deterministic rule and listed in -the scenario manifest. Each retained cell receives 10,000 planned simulation draws. The root seed is -the UTF-8 string `lane-a-e2-oc-v1`; a cell seed is derived from the root seed and canonical cell JSON -using SHA-256. Simulation order cannot affect a cell's draws. +the scenario manifest. Simulation cells use 10,000 draws unless the registered boundary rule +requires more. The root seed is the UTF-8 string `lane-a-e2-oc-v1`; cell seeds derive from the root +seed and canonical cell JSON using SHA-256. + +## Nuisance maximization and Type I control -The grid, probability model, paired-outcome construction, missingness mechanism, and candidate-rule -implementations require formal methodology review before registration. +Type I error is evaluated at the supremum over the feasible discordance nuisance domain for every +rule-relative null cell. Use analytic or exact maximization where available. Otherwise freeze the +numerical domain, tolerance, convergence tests, boundary checks, and implementation by digest. -## Outputs +For a rule whose Type I error is analytically computable, compute it exactly. The nominal level is +0.05; the `0.06` empirical ceiling is an implementation-correctness gate, not a substitute for the +rule's analytic validity proof. The substantive risk for a valid exact rule is conservativeness and +low power, which must be reported. -For every candidate rule and scenario cell, report rejection count, empirical rejection rate, Monte -Carlo standard error, and an exact binomial confidence interval. Label null cells as Type I error and -alternative cells as power. Also report refusal/`NOT_EVALUABLE` frequency, decision frequency, -class-specific error, and sensitivity to grader and apparatus violations. +Where Monte Carlo is required for admission, use simultaneous one-sided upper confidence bounds +across all registered null cells. A rule is not eligible merely because every per-cell point +estimate is below `0.06`. The simultaneous method, confidence family, multiplicity allocation, and +draw-escalation rule must be frozen before execution. Increase beyond 10,000 draws for a cell when +the registered rule places its upper bound near the admission boundary. -All cell-level records, aggregate tables, code digests, dependency lock, scenario manifest, seeds, -and report bytes must be committed as a deterministic recomputation packet by a future registered -E2 study. That packet must regenerate exactly under the registered environment. +## Missingness policies and outputs -## Candidate admission criterion +Every rule specifies one immutable missingness policy, such as dropping the pair, counting a +terminal failure against an arm, or a fully defined imputation rule. E2 stress-tests that policy +under symmetric and asymmetric mechanisms; it never chooses the most favorable policy after +simulation. -Using a one-sided nominal level of 0.05, a candidate rule is eligible for the later coupled ruling -only if the one-sided 95% exact-binomial upper confidence bound on its Type I error is at most 0.06 -in **every** registered null cell. Power has no pass threshold in this candidate; it is reported with -feasibility and practical-benefit consequences. +For every rule and cell, report the exact rejection probability or simulation rejection count/rate, +uncertainty calculation, null/alternative label, power or Type I interpretation, refusal frequency, +decision frequency, class-specific error, and sensitivity to grader/apparatus violations. -The 0.06 bound, grid, draw count, and interval choice remain formal decisions. Eligibility is not -selection: the later ruling must consider validity assumptions, power, feasibility, E4/E1 results, -and practical benefit together. +The registered packet includes all cell records, aggregate tables, code and dependency digests, +scenario manifest, seeds, and report bytes. Independent recomputation must reproduce it exactly. A +producer claim that differs from recomputation makes E2 `NOT_EVALUABLE`. -## Missing inputs and hard stops +## Hard stops and remaining freeze work -E2 is `NOT_EVALUABLE` if an accepted E4/E1 input is required but absent, a rule or grid digest does -not match registration, any registered cell is skipped, the draw count is incomplete, an output -cannot be exactly regenerated, or a producer claim differs from independent recomputation. Stop on -the first integrity mismatch and preserve the partial diagnostic separately from the registered -result. +E2 is `NOT_EVALUABLE` if a required accepted E4/E1 input is absent, a rule/grid digest differs from +registration, any cell is skipped, required draws are incomplete, the nuisance supremum is not +established, or exact recomputation fails. Partial diagnostics are retained separately and cannot +become the registered result. -Before registration, reviewers must freeze the candidate rule set, complete bounded-mean -alternative, continuous nuisance-boundary treatment, cell generator, draw count, error criterion, -compute/wall-time budget, output schema, and independent recomputation command. +Before registration, reviewers must freeze the exact conditional-rule specification, rule-relative +null partitions, nuisance maximizer, base cell generator, per-rule missingness policies, error +criterion, compute budget, output schema, and recomputation command. Held alternative rules require +their missing specifications before any implementation work begins. diff --git a/docs/preregistrations/e4-grader-validation-sap.md b/docs/preregistrations/e4-grader-validation-sap.md index 3da3bf7..1b1e44b 100644 --- a/docs/preregistrations/e4-grader-validation-sap.md +++ b/docs/preregistrations/e4-grader-validation-sap.md @@ -2,53 +2,59 @@ ## Metadata and claim boundary -- **Date:** 2026-07-17 -- **Version:** 0.1.0-candidate.1 -- **Status:** Candidate protocol — not preregistered; no observations authorized. +- **Decision date:** 2026-07-17 +- **Last updated:** 2026-07-18 +- **Version:** 0.2.0-candidate.2 +- **Status:** Candidate protocol — design decisions incorporated; not preregistered; no observations authorized. - **Study type:** Authored-fixture instrument validation; no model or provider execution. -All counts and thresholds below are **planned candidate values**. They are neither accepted design -parameters nor produced results. +Counts and thresholds below are design inputs. They are not produced results, and this SAP cannot +admit the grader without a separately registered and accepted E4 result. ## Objective -E4 tests construct validity of the deterministic editing grader against an independently authored -gold label. It does not test whether a deterministic function agrees with itself. Re-executed -score-path repeatability belongs to E1's outcome-digest checks. +E4 tests construct validity of the deterministic editing grader against independently adjudicated +human gold labels. It does not test whether a deterministic function agrees with itself. The primary question is whether machine success/failure agrees sufficiently with the adjudicated -human interpretation of the task contract to admit the instrument for calibration studies. Passing -E4 would not validate a powered comparison or establish general validity outside the frozen task -suite and grader version. - -## Planned corpus - -- 160 authored output records: 40 from each of four registered task classes. -- Within each class, 20 records are authored to satisfy the task contract and 20 are authored not to - satisfy it, for a planned 80/80 gold-label balance. -- Non-success records cover registered failure strata, including incomplete edits, out-of-scope - edits, verification failure, setup failure, timeout, and malformed output where applicable. -- Corpus records, task contracts, expected labels, failure strata, and the machine grader are frozen - by exact digest before either reviewer receives an assignment. -- Corpus construction cannot use outputs from the later E1 study or any powered campaign. - -The corpus size and strata are candidates that require methodology review before registration. +human interpretation of the task contract to admit the instrument for calibration studies. Because +the grader executes deterministically, disagreement principally diagnoses contract-to-test-system +fidelity, including ambiguous contracts, incomplete test oracles, and grader/harness +mis-specification. It is not described as random grader noise. + +## Corpus construction + +- Begin with at least 160 authored output records across four registered task classes. +- The construction target is at least 80 adjudicated-positive and 80 adjudicated-negative records. +- Each class includes satisfying, clearly non-satisfying, boundary, and near-miss records. +- Boundary categories, minimum effective adjudicated counts per stratum, deterministic top-up + rules, and the ambiguity threshold must be fixed before authoring starts. +- A top-up appends a new retained record under the predeclared rule; it never replaces or suppresses + a reviewed record. The final registered corpus may therefore exceed 160 records. +- Corpus records, task contracts, author-intended labels, failure/boundary strata, and the machine + grader are frozen by exact digest before reviewer assignment. +- Construction cannot use outputs from E1 or a powered campaign. + +The corpus is a versioned, extensible validation instrument. Suite growth requires a new corpus +version and revalidation rather than silently carrying forward the prior admission. + +Registration remains blocked until the boundary taxonomy, per-stratum minimums, top-up rule, and +ambiguity threshold are populated with exact values and reviewed. ## Independent labels and adjudication Two qualified reviewers independently label every record without seeing the machine grade, the -other reviewer's label, or a success/non-success target. Reviewer order is randomized. Reviewers +other reviewer's label, or the author-intended target. Reviewer order is randomized. Reviewers record a binary label, rationale code, and ambiguity flag. -Agreement between reviewers becomes the provisional gold label. A disagreement or any ambiguity -flag goes to a third adjudicator, who sees the task contract and both rationales but remains blinded -to the machine grade. The adjudicated label is the analysis gold label. Reviewer identities, -qualification criteria, conflicts, assignments, and adjudications must be retained in the private -study record; only a sanitized role-level report is public. +Agreement becomes the provisional gold label. A disagreement or any ambiguity flag goes to a third +adjudicator, who sees the task contract and both rationales but remains blinded to the machine +grade. The adjudicated label is the analysis gold label. The author-intent-versus-adjudicated-gold +disagreement rate and ambiguity rate are reported as contract-ambiguity diagnostics. -## Analysis +## Analysis and certification granularity -Using the adjudicated human label as truth and machine success as the prediction, report `TP`, `TN`, +Using adjudicated human label as truth and machine success as the prediction, report `TP`, `TN`, `FP`, and `FN`, then compute: \[ @@ -56,28 +62,34 @@ Using the adjudicated human label as truth and machine success as the prediction \mathrm{specificity}=\frac{TN}{TN+FP}. \] -Also report positive predictive value, negative predictive value, class-stratified confusion -matrices, failure-stratum errors, raw reviewer agreement, and Cohen's kappa. Sensitivity and -specificity receive two-sided 95% Wilson intervals. Reviewer-agreement measures are disclosures, -not substitute admission criteria. +Certification applies to the overall adjudicated-positive population and the overall +adjudicated-negative population. Class- and stratum-specific confusion matrices are diagnostics; +their smaller denominators do not independently carry the overall Wilson-bound admission rule. + +Also report positive and negative predictive values, class/stratum confusion matrices, raw reviewer +agreement, Cohen's kappa, author/adjudicator disagreement, and ambiguity. Sensitivity and +specificity receive two-sided 95% Wilson intervals. -## Candidate admission rule +## Admission rule -The instrument is only a **candidate for calibration admission** when all of these planned criteria -hold: +The instrument is eligible for explicit calibration admission only when all of these criteria hold: - point sensitivity is at least 0.90; - point specificity is at least 0.95; -- the two-sided 95% Wilson lower bound is at least 0.80 for both sensitivity and specificity; -- all 160 records have complete independent labels and any required adjudication; and +- the two-sided 95% Wilson lower bound is at least 0.80 for both overall sensitivity and overall + specificity; +- the frozen minimum effective adjudicated counts and every corpus-stratum requirement are met; +- every record has complete independent labels and any required adjudication; and - no corpus, grader, task-contract, blinding, or byte-binding violation occurred. -These numerical thresholds remain decisions for formal review. Even if achieved, the registered -E4 result must still be accepted explicitly; this draft cannot admit the instrument by itself. +The higher specificity threshold deliberately prioritizes false-pass avoidance: crediting a failure +as success can inflate a later candidate score, while a false fail is conservative. Passing this +rule does not establish validity outside the frozen suite and grader version. ## Missingness and hard stops -Any missing label, lost rationale, failed blinding check, post-assignment corpus change, or absent -digest makes E4 `NOT_EVALUABLE`. Records are not replaced after review begins. E4 stops immediately -on unblinding or evidence-integrity failure. The review schedule, reviewer-hour budget, conflict -rule, and exact public report schema must be frozen before registration. +Any missing label, lost rationale, failed blinding check, post-assignment mutation, absent digest, +or unregistered top-up makes E4 `NOT_EVALUABLE`. E4 stops immediately on unblinding or +evidence-integrity failure. Before registration, reviewers must freeze reviewer qualifications, +conflict rules, assignments, hour budget, corpus-stratum values, evidence schema, and exact public +report schema. diff --git a/tests/test_calibration_saps.py b/tests/test_calibration_saps.py index 8a26816..1d877a4 100644 --- a/tests/test_calibration_saps.py +++ b/tests/test_calibration_saps.py @@ -15,43 +15,78 @@ SAP_DIR / "e1-lane-a-aa-sap.md", SAP_DIR / "e2-operating-characteristics-sap.md", ) +STATUS = ( + "Status:** Candidate protocol — design decisions incorporated; " + "not preregistered; no observations authorized." +) class CalibrationSapTests(unittest.TestCase): - def test_saps_are_explicitly_unregistered_and_public_safe(self) -> None: + def test_saps_are_unregistered_non_authorizing_and_public_safe(self) -> None: for path in SAP_FILES: with self.subTest(path=path.name): text = path.read_text(encoding="utf-8") - self.assertIn( - "Status:** Candidate protocol — not preregistered; no observations authorized.", - text, - ) - self.assertNotIn("Status:** Preregistered Draft", text) + self.assertIn(STATUS, text) + self.assertNotIn("Status:** Preregistered", text) assert_publication_content_is_safe(text) - def test_e1_uses_m_one_and_fixed_terminal_counts(self) -> None: + def test_umbrella_declares_dependency_and_opaque_profile(self) -> None: + text = (SAP_DIR / "calibration-sap.md").read_text(encoding="utf-8") + normalized = " ".join(text.split()) + self.assertIn("E4 and E1 may proceed independently", text) + self.assertIn("base operating-characteristics grid", text) + self.assertIn("opaque registered input", text) + self.assertIn("before the E4/E1 results are observed", normalized) + + def test_e1_localizes_exact_identity_and_uses_session_pairs(self) -> None: text = (SAP_DIR / "e1-lane-a-aa-sap.md").read_text(encoding="utf-8") self.assertIn("`m = 1`", text) self.assertIn("120 sessions total", text) + self.assertIn("120 primary cross-arm comparisons", text) self.assertIn("240 terminal arm attempts total", text) self.assertIn("actual chronological session position", text) - self.assertIn("`INVARIANCE_NOT_REFUTED`", text) - self.assertIn("`INVARIANCE_REFUTED`", text) - self.assertIn("`APPARATUS_NOT_ADMISSIBLE`", text) - self.assertIn("`NOT_EVALUABLE`", text) - self.assertNotIn("objective is to estimate", text.lower()) - self.assertNotIn("pearson correlation", text.lower()) - - def test_e4_and_e2_numbers_are_labeled_candidate(self) -> None: - e4 = (SAP_DIR / "e4-grader-validation-sap.md").read_text(encoding="utf-8") - e2 = (SAP_DIR / "e2-operating-characteristics-sap.md").read_text( + self.assertIn("`ENGINE_STAGE_DIVERGENCE_UNDER_REGISTERED_APPARATUS`", text) + self.assertIn("authoritative, request-bound evidence", text) + self.assertIn("Timing is corroborative only", text) + self.assertIn("approximately `3 / 120 = 2.5%`", text) + self.assertIn("repeated-task dependence", text) + self.assertNotIn("engine_nondeterminism", text) + + def test_e1_conservative_state_precedence_is_explicit(self) -> None: + text = (SAP_DIR / "e1-lane-a-aa-sap.md").read_text(encoding="utf-8") + states = [ + "1. `NOT_EVALUABLE`", + "2. `APPARATUS_NOT_ADMISSIBLE`", + "3. `INVARIANCE_REFUTED`", + "4. `INVARIANCE_NOT_REFUTED`", + ] + positions = [text.index(state) for state in states] + self.assertEqual(positions, sorted(positions)) + self.assertIn("including a failure that co-occurs with divergence", text) + + def test_e4_certifies_overall_gold_populations_and_freezes_strata(self) -> None: + text = (SAP_DIR / "e4-grader-validation-sap.md").read_text(encoding="utf-8") + self.assertIn("at least 160 authored output records", text) + self.assertIn("at least 80 adjudicated-positive", text) + self.assertIn("boundary, and near-miss records", text) + self.assertIn("deterministic top-up", text) + self.assertIn("overall adjudicated-positive population", text) + self.assertIn("point sensitivity is at least 0.90", text) + self.assertIn("point specificity is at least 0.95", text) + self.assertIn("false-pass avoidance", text) + + def test_e2_uses_rule_relative_exact_or_simultaneous_admission(self) -> None: + text = (SAP_DIR / "e2-operating-characteristics-sap.md").read_text( encoding="utf-8" ) - self.assertIn("160 authored output records", e4) - self.assertIn("planned candidate values", e4) - self.assertIn("10,000 planned simulation draws", e2) - self.assertIn("planned candidate value", e2) - self.assertIn("must be committed", e2) + self.assertIn("exact conditional paired binary", text) + self.assertIn("remain held design slots", text) + self.assertIn("rule-relative null cell", text) + self.assertIn("supremum over the feasible discordance nuisance domain", text) + self.assertIn("simultaneous one-sided upper confidence bounds", text) + self.assertIn("10,000 draws", text) + self.assertIn("`N = 60` is a hypothetical feasibility point", text) + self.assertIn("producer claim that differs from recomputation", text) if __name__ == "__main__":