diff --git a/README.md b/README.md index ae000a7..1b616b9 100644 --- a/README.md +++ b/README.md @@ -261,6 +261,20 @@ fixture inputs, not ACE outputs. This is exact forecast/result scoring, not a hi contemporaneous forecast, model-skill finding, empirical calibration curve, population reliability estimate, causal claim, or human-benefit finding. +P2C10 independently reproduces correction-quality measurement over the U.S. Bureau of Labor +Statistics public errata family. The recorded July 1, 2025 JOLTS release `USDL-25-1087` and BLS's +July 2 correction are admitted as distinct immutable Observations. Treatment and control preserve +both source identities, the exact correction link, two reviewed Actions each, and the same matched +conditions; treatment renders the required minus sign in `−39,000`, while control retains the +reported pre-correction form. Exact reviews score `1.0` and `0.0`, so the unchanged domain-neutral +contract classifies the bounded two-pair difference as useful and emits only a non-effective +proposal. + +The historical wrong form is derived from BLS's explicit statement that the sentence required a +missing minus sign; the currently archived release is already corrected. This is recorded replay +over one correction, not live monitoring, statistical validation of JOLTS, population correction +performance, causality, or human benefit. No Domain Pack, connector, or Core contract changes. + ## What the public World proof demonstrates Generate a self-contained visual Reality Brief and its exact machine-readable backing data: @@ -361,6 +375,7 @@ $PY -m scripts.p2c6_contradiction_attention_outcome "$WORKSPACE" $PY -m scripts.p2c7_correction_detection_delay_outcome "$WORKSPACE" $PY -m scripts.p2c8_correction_revision_stability_outcome "$WORKSPACE" $PY -m scripts.p2c9_forecast_calibration_outcome "$WORKSPACE" +$PY -m scripts.p2c10_independent_correction_reproduction "$WORKSPACE" ``` The released 0.9.0 gates are reproducible through the locked environment, as CI does. The @@ -397,8 +412,8 @@ The complete governed product-journey evidence is recorded in [`docs/audits/world-intelligence-p2c2-governed-reality-brief-2026-08-10.md`](docs/audits/world-intelligence-p2c2-governed-reality-brief-2026-08-10.md). The source-checkout measured-feedback candidate is recorded in [`docs/audits/world-intelligence-p2c3-measured-feedback-2026-08-10.md`](docs/audits/world-intelligence-p2c3-measured-feedback-2026-08-10.md). -The latest withheld-result forecast-scoring candidate is recorded in -[`docs/audits/world-intelligence-p2c9-forecast-calibration-outcome-2026-08-10.md`](docs/audits/world-intelligence-p2c9-forecast-calibration-outcome-2026-08-10.md). +The latest independent-source correction candidate is recorded in +[`docs/audits/world-intelligence-p2c10-independent-correction-reproduction-2026-08-10.md`](docs/audits/world-intelligence-p2c10-independent-correction-reproduction-2026-08-10.md). Release-level scope and evidence are recorded in [`docs/releases/world-intelligence-p2c2-v0.9.0.md`](docs/releases/world-intelligence-p2c2-v0.9.0.md), @@ -435,8 +450,9 @@ detection delay against a product target. P2C8 measures exact unaffected-claim i preservation across real Brief contracts while holding correction semantics, source coverage, and claim count constant. P2C9 adds an exact withheld-result forecast record and derives a single-event Brier contribution while explicitly withholding any population-calibration or model-skill claim. -The next bounded measurement work is another independently sourced correction event or independent -Market reproduction. +P2C10 reproduces correction-quality measurement over an independently sourced BLS erratum without +changing the shared contracts. The next bounded measurement work is independent Market +reproduction. Separately reviewed opt-in network transport and P2D multi-source conflict/correction with LIVE inputs remain independent work. None of these steps may add autonomous publishing, delivery, persuasion, or action authority to a Domain Pack. diff --git a/ROADMAP.md b/ROADMAP.md index 673fb8c..2b1d209 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -16,7 +16,7 @@ source code into the platform. See the [0.9.0 release record](docs/releases/world-intelligence-p2c2-v0.9.0.md). -## Candidate — P2C3–P2C9 measured feedback and product-owned outcomes +## Candidate — P2C3–P2C10 measured feedback and product-owned outcomes - The exact 0.9.0 Brief and a World-owned source-only control pass through two matched reviewed export pairs under one frozen structural citation-coverage criterion. @@ -51,6 +51,11 @@ See the [0.9.0 release record](docs/releases/world-intelligence-p2c2-v0.9.0.md). resolves the binary event; exact reviews derive single-event Brier quality of `0.9375` and `0.4375`. This proves forecast/result scoring and leakage-resistant record order, not historical contemporaneity, probability generation by ACE, model skill, or population calibration. +- A seventh frozen criterion repeats correction-quality measurement over the independent BLS public + errata source family. Treatment and control retain the same release, correction, correction link, + source coverage, reviewed workflow, and matched conditions; only treatment renders the exact + corrected minus sign. The unchanged shared contract derives `1.0` versus `0.0`, classifies the + bounded two-pair difference useful, and leaves the proposal non-effective and unapplied. - This is source-checkout evidence against stacked Core candidates, not a released World capability, human benefit finding, causal claim, network-freshness proof, or applied governance change. @@ -72,6 +77,10 @@ performance, or a general semantic-equivalence engine. The stacked [P2C9 work packet](docs/design/world-intelligence-p2c9-forecast-calibration-outcome-work-packet-v1.md) freezes an exact forecast/result scoring boundary without claiming a historical forecast, model skill, or population calibration. +The stacked +[P2C10 work packet](docs/design/world-intelligence-p2c10-independent-correction-reproduction-work-packet-v1.md) +freezes independent BLS correction reproduction without claiming live monitoring, statistical +validity, population correction performance, causality, or human benefit. ## Next — trustworthy live orientation diff --git a/docs/audits/world-intelligence-p2c10-independent-correction-reproduction-2026-08-10.md b/docs/audits/world-intelligence-p2c10-independent-correction-reproduction-2026-08-10.md new file mode 100644 index 0000000..c483125 --- /dev/null +++ b/docs/audits/world-intelligence-p2c10-independent-correction-reproduction-2026-08-10.md @@ -0,0 +1,137 @@ +# World Intelligence P2C10 independent correction reproduction audit — 2026-08-10 + +Status: **stacked candidate evidence only; not a release, SI4 pass, live-monitoring result, +population-performance finding, or applied governance change** + +## Source identity + +- World base: P2C9 commit `53feadb40fcc93d23f326b16979ed6640471c4cf` +- World branch: `codex/world-independent-correction-reproduction` +- Core dependency: exact observed-result candidate commit + `433e3d16c5458c975557dcd1552824fb959d4d12` +- Released identity intentionally unchanged: `ace-domain-world-intelligence==0.9.0`, + `ace-core>=0.5.0,<0.6` + +## Exact point-in-time result + +The frozen BLS fixture is +`sha256:981183a2464f74f4421bd0a6470f0342a5abae8cfebc1b5a8da1562df26babb1`. +It binds release `USDL-25-1087` to correction `bls-errata-2025-07-01-jolts`, with recorded +Observation identities: + +```text +original: observation:faf25d26cc88802368cabf3e17538a7d +correction: observation:3a1351d6ac306374b8a5b472c192d2b9 +``` + +The point-run treatment artifact was +`official_correction_artifact:34d538fbd2c3a6ef6d3e79d0fdb8a344` with material +`sha256:34d538fbd2c3a6ef6d3e79d0fdb8a3440c06b02a1e156a6001bfc1491fd86e21`. +The matched control was +`official_correction_artifact:ab6cfedd1b76602e51c0c8d3c1dbd6d2` with material +`sha256:ab6cfedd1b76602e51c0c8d3c1dbd6d22e5c76b01b1eb59350a8e1bdaa7d8207`. + +Both artifacts name the same exact release and correction Observations, `corrects` relation, +source-policy digest, and reviewed workflow. Treatment renders `−39,000`; control renders the +reported pre-correction `(39,000)` form. All four reviews independently recorded complete source +coverage, visible correction linkage, and preserved prior history: + +```text +treatment scores: 1.0, 1.0 +control scores: 0.0, 0.0 +matched pairs: 2 +mean effect: 1.0 +classification: useful +proposal action: promote +proposal live effect: false +historical replay: true +replay reauthorization: false +``` + +The exact point-run evaluation was `impact_evaluation:5490b5c680940d88d25d9dacca67103b` +with material +`sha256:5490b5c680940d88d25d9dacca67103b6fb33b8d8c98e0e756488f45dece964e`. +The exact non-effective proposal was +`impact_governance_proposal:da09bf45bb9f48d6cefbb3fb49122e67` with material +`sha256:da09bf45bb9f48d6cefbb3fb49122e679d010cd39140950a57c447e2465f7f86`. + +Point-run artifact and evaluation identities include exact record-availability and reviewed-action +times. Historical replay in the same durable store returns those exact identities without +reauthorization. Fresh hosts reproduce the fixture digest, source content identities, scores, +classification, and proposal semantics without pretending independent wall-clock availability +coordinates are identical. + +## Verification + +The frozen World dependency versions plus the stacked Core candidate and separately packaged +reference action adapter produced: + +```text +python -B -m pytest domain_packs/tests/test_p2c3_measured_feedback.py \ + domain_packs/tests/test_p2c4_reviewed_impact_disposition.py \ + domain_packs/tests/test_p2c5_citation_correctness_outcome.py \ + domain_packs/tests/test_p2c6_contradiction_attention_outcome.py \ + domain_packs/tests/test_p2c7_correction_detection_delay_outcome.py \ + domain_packs/tests/test_p2c8_correction_revision_stability_outcome.py \ + domain_packs/tests/test_p2c9_forecast_calibration_outcome.py \ + domain_packs/tests/test_p2c10_independent_correction_reproduction.py -q --tb=short +30 passed in 11.01s + +python -B -m pytest -q --tb=short +113 passed in 24.20s + +python -B -m pytest adapters/federal_register_source/tests -q --tb=short +26 passed in 0.25s + +python -B -m pytest tests/test_release_contract.py -q --tb=short +7 passed in 0.05s + +# Installed public ace-core==0.5.0; candidate tests skip explicitly. +python -B -m pytest -q --tb=short -rs +83 passed, 30 skipped in 14.10s + +ruff check --no-cache +PASS + +ruff format --check --no-cache +3 files already formatted + +UV_CACHE_DIR=/tmp/ace-p2c10-uv-cache uv build --out-dir /tmp/ace-p2c10-dist-20260810 +Successfully built unchanged 0.9.0 source distribution and inert data-only wheel + +git diff --check +PASS +``` + +The thirty public-Core skips are explicit candidate boundaries: P2C3 through P2C10 require +unreleased stacked Core measured-impact contracts. The public P2C2 journey and every released +boundary remain green. The wheel contains 45 inert Domain Pack JSON/metadata files, no Python or +entry points, and retains `ace-core>=0.5.0,<0.6`. + +Repository-wide Ruff remains an inherited release-hygiene blocker. With the same locked Ruff +version, both the P2C9 parent and this P2C10 worktree report exactly 14 lint findings and 12 format +targets. Scoped P2C10 checks and `git diff --check` are green. This packet does not rewrite +unrelated history, but release closeout must reconcile the repository-wide gate before publication. + +## Claim boundary + +BLS's public errata states that the sentence required a missing minus sign; the current archived +release already exposes the corrected form. The fixture explicitly derives the historical original +form from that erratum. It does not claim to preserve original response bytes. + +The World review exact-loads the artifact and both immutable source Observations, derives source +coverage, correction linkage, prior-record preservation, corrected-form equality, and stale-form +absence, then records an exact observed result. Core and Intelligence receive only domain-neutral +records, conditions, scores, and classification. Historical replay requires no new authority, and +the proposal is non-effective, non-selectable, unapplied, and subject to separate human review. + +This is one hermetic recorded BLS correction. It is not live monitoring, network-arrival evidence, +statistical validation of JOLTS, population correction performance, a general source-independence +claim, causality, general Brief quality, or human benefit. + +## Remaining work + +Independent Market reproduction, combined-main review/CI, public Core artifacts, +repository-wide lint/format reconciliation, security/release checks, and opt-in live transport +remain future bounded work. Core issue #49 F1, F3, and F5 still require explicit 0.6 release-owner +disposition; this World packet neither implements nor re-dates them. diff --git a/docs/design/world-intelligence-p2c10-independent-correction-reproduction-work-packet-v1.md b/docs/design/world-intelligence-p2c10-independent-correction-reproduction-work-packet-v1.md new file mode 100644 index 0000000..362fc53 --- /dev/null +++ b/docs/design/world-intelligence-p2c10-independent-correction-reproduction-work-packet-v1.md @@ -0,0 +1,117 @@ +# World Intelligence P2C10 independent correction reproduction work packet (v1) + +**Status:** stacked source-checkout candidate; this packet does not release World Intelligence, +apply a governance proposal, close ACE Core issue #38, pass SI4, or complete ACE 0.6.0. + +**Frozen:** 2026-08-10 from World P2C9 commit +`53feadb40fcc93d23f326b16979ed6640471c4cf`, stacked on the Core exact observed-result +provenance candidate `433e3d16c5458c975557dcd1552824fb959d4d12`. + +## Objective + +Falsify source-family coupling by reproducing the measured correction journey over a materially +different real publisher and source policy while leaving the shared Core + Intelligence contracts +unchanged: + +```text +BLS release Observation + later BLS erratum Observation + -> treatment corrected artifact / stale-form control + -> Decision -> reviewed Action -> independent review -> exact Outcome + -> useful / harmful / unproven evaluation -> proposal only +``` + +The source pair is the BLS Job Openings and Labor Turnover release `USDL-25-1087` and its public +erratum. The [archived release](https://www.bls.gov/news.release/archives/jolts_07012025.htm) now +contains the corrected `−39,000` form. The [BLS errata page](https://www.bls.gov/errata/) states +that the July 1, 2025 sentence required a missing minus sign and that corrections were made July 2. +The fixture therefore labels the pre-correction form as derived from that explicit erratum; it does +not pretend the current archive still exposes the superseded bytes. + +## Frozen product policy and matched control + +World owns `world_independent_official_correction_statement_quality` version `candidate-1`. A +review scores `1.0` only when all of the following are exact and inspectable: + +1. both the original-release and correction Observations are present; +2. the correction names the exact release it corrects; +3. the original immutable record remains loadable and unchanged; +4. the rendered statement equals the public corrected statement; and +5. the reported pre-correction form is absent. + +Otherwise the score is `0.0`. Treatment and control share both source references, correction +linkage, policy digest, two reviewed Action workflows, observation window, and matched conditions. +Treatment renders `The number of job openings decreased in federal government (−39,000).`; control +retains the missing-sign form `(39,000)`. Source coverage and action volume therefore cannot explain +the score difference. + +The criterion requires two matched pairs and an effect of `1.0`. Intelligence derives useful, +harmful, or unproven under that frozen rule. Core appends exact reviews, Outcomes, evaluation, and +proposal history. A useful result maps only to a non-effective, non-selectable `promote` proposal +requiring separate human review. + +## Exact acceptance + +P2C10 must: + +1. rerun P2C2 through P2C9 and preserve all prior immutable results and proposals; +2. admit the recorded BLS release and erratum as distinct exact Observations without network use; +3. retain BLS source vocabulary and historical-original derivation policy only in World; +4. append treatment and stale-form control artifacts that name the same exact source pair and + correction relation; +5. complete two reviewed treatment Actions and two reviewed control Actions; +6. exact-load each artifact and both source Observations before deriving the product score; +7. append four authenticated reviews and four Core Outcomes naming those reviews as observed + results; +8. classify `1.0` versus `0.0` over two pairs as useful and emit proposal-only `promote`; +9. replay without reauthorization and reproduce the fixture digest, scores, classification, and + proposal semantics across fresh hosts; and +10. reject drifted source linkage, duplicate source identity, changed statement forms, invented + scores, and fixture-policy drift. + +Stacked Core tests remain authoritative for missing attribution, condition mismatch, cutoff +leakage, unavailable Outcomes, duplicate/replayed evidence, interruption, restart, and denied +authority. This packet exercises those unchanged contracts through a second real source family +rather than duplicating their lower-level tests in World. + +## Ownership boundary + +World owns BLS vocabulary, source URLs, the historical-original derivation disclosure, correction +statement policy, fixture, matched control, review contract, and product evidence. Core owns durable +identities, append-only records, provenance, Decisions, reviewed Actions, Outcomes, authority, and +replay. Intelligence owns domain-neutral conditions, matched evaluation, uncertainty, +classification, and proposal contracts. No BLS, JOLTS, errata, or minus-sign noun moves into Core or +Intelligence. + +The installable Domain Pack remains inert JSON and unchanged. The Federal Register connector is +unchanged and is not used to fetch BLS. The fixture is recorded, hermetic, and network-free. + +## Files, rollback, and deletion criteria + +This packet owns: + +- `domain_packs/tests/fixtures/p2c10_bls_correction_pair.json`; +- `scripts/p2c10_independent_correction_reproduction.py`; +- `domain_packs/tests/test_p2c10_independent_correction_reproduction.py`; +- the additive P2C9 state handoff; and +- this work packet, its audit, and restrained README/roadmap references. + +It changes no shipped Domain Pack, connector, package version, dependency range, lockfile, release +record, Core contract, or public artifact. Rollback removes the harness, tests, handoff, fixture, +and candidate documentation. Records already persisted by a host remain immutable history. + +Delete or replace this fixture only if BLS removes the public correction evidence, the exact URLs +cannot be independently verified, or a stronger redistributable snapshot supersedes it. Such a +change requires a new fixture identity and cannot rewrite prior evidence. + +## Non-claims and next packet + +This packet does not establish live monitoring, network freshness, the statistical validity of the +JOLTS estimate, general correction quality, population performance, source independence beyond the +two exercised families, causality, general Brief quality, or human benefit. It does not apply a +proposal or grant authority to a Domain Pack. + +The next bounded falsification packet is an independent Market Intelligence reproduction of the +same public Core + Intelligence contracts. Combined-main review/CI, public Core artifacts, +compatibility and security checks, repository-wide hygiene, and separately reviewed opt-in live +transport remain release work. Core issue #49 F1, F3, and F5 still require explicit 0.6 +release-owner disposition; this packet does not implement, defer, or re-date them. diff --git a/domain_packs/tests/fixtures/p2c10_bls_correction_pair.json b/domain_packs/tests/fixtures/p2c10_bls_correction_pair.json new file mode 100644 index 0000000..ab1e3e1 --- /dev/null +++ b/domain_packs/tests/fixtures/p2c10_bls_correction_pair.json @@ -0,0 +1,30 @@ +{ + "fixture_id": "bls-jolts-may-2025-minus-sign-correction", + "network_access": false, + "recorded_at": "2026-08-10T23:50:00Z", + "source_policy": { + "publisher": "U.S. Bureau of Labor Statistics", + "source_family": "bls_public_errata", + "release_uri": "https://www.bls.gov/news.release/archives/jolts_07012025.htm", + "errata_uri": "https://www.bls.gov/errata/", + "recorded_replay": true, + "historical_original_form_derived_from_erratum": true, + "statistical_validity_claimed": false + }, + "original": { + "record_id": "USDL-25-1087", + "release_date": "2025-07-01", + "published_at": "2025-07-01T14:00:00Z", + "release_uri": "https://www.bls.gov/news.release/archives/jolts_07012025.htm", + "reported_sentence_without_required_minus_sign": "The number of job openings decreased in federal government (39,000)." + }, + "correction": { + "record_id": "bls-errata-2025-07-01-jolts", + "corrects_record_id": "USDL-25-1087", + "date_added": "2025-07-01", + "corrected_at": "2025-07-02T14:00:00Z", + "errata_uri": "https://www.bls.gov/errata/", + "correction_description": "The following sentence requires correction to add a missing minus sign.", + "corrected_sentence": "The number of job openings decreased in federal government (−39,000)." + } +} diff --git a/domain_packs/tests/test_p2c10_independent_correction_reproduction.py b/domain_packs/tests/test_p2c10_independent_correction_reproduction.py new file mode 100644 index 0000000..795d5c0 --- /dev/null +++ b/domain_packs/tests/test_p2c10_independent_correction_reproduction.py @@ -0,0 +1,170 @@ +from __future__ import annotations + +import copy +import importlib.util +import json + +import pytest +from pydantic import ValidationError + + +def _require_candidate_contracts() -> None: + if importlib.util.find_spec("ace.application.measured_impact") is None: + pytest.skip("P2C10 requires the stacked ACE Core measured-impact candidate") + from ace.intelligence import ImpactOutcomeMeasuresV1Alpha1 + + if "observed_result" not in ImpactOutcomeMeasuresV1Alpha1.model_fields: + pytest.skip("P2C10 requires exact observed-result provenance from the stacked Core candidate") + if importlib.util.find_spec("ace_reference_workspace_action") is None: + pytest.skip("P2C10 requires the separately packaged Core reference adapter") + + +def test_recorded_bls_fixture_freezes_an_independent_exact_correction_pair() -> None: + _require_candidate_contracts() + from scripts.p2c10_independent_correction_reproduction import ( + bls_correction_fixture_digest, + load_bls_correction_fixture, + ) + + fixture = load_bls_correction_fixture() + + assert fixture["network_access"] is False + assert fixture["source_policy"]["source_family"] == "bls_public_errata" + assert fixture["source_policy"]["historical_original_form_derived_from_erratum"] is True + assert fixture["original"]["record_id"] == "USDL-25-1087" + assert fixture["correction"]["corrects_record_id"] == fixture["original"]["record_id"] + assert fixture["original"]["reported_sentence_without_required_minus_sign"].endswith("(39,000).") + assert fixture["correction"]["corrected_sentence"].endswith("(−39,000).") + assert fixture["original"]["release_uri"].startswith("https://www.bls.gov/") + assert fixture["correction"]["errata_uri"] == "https://www.bls.gov/errata/" + assert ( + bls_correction_fixture_digest(fixture) + == "sha256:981183a2464f74f4421bd0a6470f0342a5abae8cfebc1b5a8da1562df26babb1" + ) + + +@pytest.mark.asyncio +async def test_independent_source_correction_becomes_an_exact_measured_outcome(tmp_path) -> None: + _require_candidate_contracts() + from scripts.p2c10_independent_correction_reproduction import run_independent_correction_reproduction + + result = await run_independent_correction_reproduction(tmp_path) + + assert result["review_policy"]["reviewer_ref"] == "principal:world-independent-correction-reviewer" + assert result["evaluation"]["criterion"]["requires_observed_result"] is True + assert result["evaluation"]["classification"] == "useful" + assert result["evaluation"]["matched_pair_count"] == 2 + assert result["evaluation"]["mean_effect"] == 1.0 + assert result["proposal"]["action"] == "promote" + assert result["proposal"]["live_effect"] is False + assert result["proposal"]["selectable"] is False + assert result["proposal"]["requires_human_review"] is True + assert result["replay"]["historical"] is True + assert result["replay"]["no_reauthorization"] is True + assert result["scope"]["independent_source_family_reproduction"] is True + assert result["scope"]["domain_neutral_core_contract_unchanged"] is True + + +@pytest.mark.asyncio +async def test_source_coverage_linkage_and_reviewed_workflow_are_matched(tmp_path) -> None: + _require_candidate_contracts() + from scripts.p2c10_independent_correction_reproduction import run_independent_correction_reproduction + + result = await run_independent_correction_reproduction(tmp_path) + treatment_artifact = result["artifacts"]["treatment"] + control_artifact = result["artifacts"]["control"] + treatment_reviews = result["observed_results"]["treatment"] + control_reviews = result["observed_results"]["control"] + + assert treatment_artifact["original_observation"] == control_artifact["original_observation"] + assert treatment_artifact["correction_observation"] == control_artifact["correction_observation"] + assert treatment_artifact["corrects_source_record_id"] == control_artifact["corrects_source_record_id"] + for reviews in (treatment_reviews, control_reviews): + assert {item["source_coverage_complete"] for item in reviews} == {True} + assert {item["correction_link_visible"] for item in reviews} == {True} + assert {item["prior_record_preserved"] for item in reviews} == {True} + assert {item["corrected_statement_exact"] for item in treatment_reviews} == {True} + assert {item["stale_form_present"] for item in treatment_reviews} == {False} + assert {item["correction_quality_score"] for item in treatment_reviews} == {1.0} + assert {item["corrected_statement_exact"] for item in control_reviews} == {False} + assert {item["stale_form_present"] for item in control_reviews} == {True} + assert {item["correction_quality_score"] for item in control_reviews} == {0.0} + assert result["scope"]["prior_record_preserved"] is True + assert result["scope"]["network_access"] is False + assert result["scope"]["population_correction_performance_claimed"] is False + + +@pytest.mark.asyncio +async def test_drifted_source_link_duplicate_identity_and_invented_score_fail_closed(tmp_path) -> None: + _require_candidate_contracts() + from scripts.p2c10_independent_correction_reproduction import ( + IndependentCorrectionReviewV1Alpha1, + OfficialCorrectionArtifactV1Alpha1, + load_bls_correction_fixture, + run_independent_correction_reproduction, + validate_bls_correction_fixture, + ) + + fixture = copy.deepcopy(load_bls_correction_fixture()) + fixture["correction"]["corrects_record_id"] = "USDL-OTHER" + with pytest.raises(AssertionError, match="correction linkage changed"): + validate_bls_correction_fixture(fixture) + + fixture = copy.deepcopy(load_bls_correction_fixture()) + fixture["source_policy"]["statistical_validity_claimed"] = True + with pytest.raises(AssertionError, match="source policy changed"): + validate_bls_correction_fixture(fixture) + + result = await run_independent_correction_reproduction(tmp_path) + artifact = copy.deepcopy(result["artifacts"]["treatment"]) + artifact["corrects_source_record_id"] = "USDL-OTHER" + artifact["artifact_id"] = None + artifact["artifact_digest"] = None + with pytest.raises(ValidationError, match="lost its exact correction linkage"): + OfficialCorrectionArtifactV1Alpha1.model_validate_json(json.dumps(artifact)) + + artifact = copy.deepcopy(result["artifacts"]["treatment"]) + artifact["correction_observation"] = artifact["original_observation"] + artifact["artifact_id"] = None + artifact["artifact_digest"] = None + with pytest.raises(ValidationError, match="requires distinct exact source records"): + OfficialCorrectionArtifactV1Alpha1.model_validate_json(json.dumps(artifact)) + + artifact = copy.deepcopy(result["artifacts"]["treatment"]) + artifact["displayed_statement"] = "The estimate was corrected." + artifact["artifact_id"] = None + artifact["artifact_digest"] = None + with pytest.raises(ValidationError, match="introduced an unreviewed statement form"): + OfficialCorrectionArtifactV1Alpha1.model_validate_json(json.dumps(artifact)) + + review = copy.deepcopy(result["observed_results"]["control"][0]) + review["correction_quality_score"] = 1.0 + review["review_id"] = None + review["review_digest"] = None + with pytest.raises(ValidationError, match="score differs from the frozen product rule"): + IndependentCorrectionReviewV1Alpha1.model_validate_json(json.dumps(review)) + + +@pytest.mark.asyncio +async def test_independent_reproduction_is_substantively_deterministic_across_fresh_hosts(tmp_path) -> None: + _require_candidate_contracts() + from scripts.p2c10_independent_correction_reproduction import run_independent_correction_reproduction + + first_root = tmp_path / "first" + second_root = tmp_path / "second" + first_root.mkdir() + second_root.mkdir() + first = await run_independent_correction_reproduction(first_root) + second = await run_independent_correction_reproduction(second_root) + + assert first["source_pair"]["fixture_digest"] == second["source_pair"]["fixture_digest"] + assert first["review_policy"]["policy_id"] == second["review_policy"]["policy_id"] + assert first["review_policy"]["policy_version"] == second["review_policy"]["policy_version"] + for variant in ("treatment", "control"): + first_scores = [item["correction_quality_score"] for item in first["observed_results"][variant]] + second_scores = [item["correction_quality_score"] for item in second["observed_results"][variant]] + assert first_scores == second_scores + for field in ("classification", "matched_pair_count", "mean_effect"): + assert first["evaluation"][field] == second["evaluation"][field] + for field in ("action", "live_effect", "selectable", "requires_human_review"): + assert first["proposal"][field] == second["proposal"][field] diff --git a/scripts/p2c10_independent_correction_reproduction.py b/scripts/p2c10_independent_correction_reproduction.py new file mode 100644 index 0000000..8e8a7de --- /dev/null +++ b/scripts/p2c10_independent_correction_reproduction.py @@ -0,0 +1,852 @@ +"""Reproduce measured correction quality over an independent BLS source family.""" + +from __future__ import annotations + +import asyncio +import json +from datetime import datetime +from pathlib import Path +from typing import Any, Literal, Self + +from ace.application import MeasuredImpactService +from ace.core import ( + AuthenticatedRuntimeContextV1Alpha1, + CapabilityArtifactIdentityV1Alpha1, + GovernedOperationBindingV1Alpha1, + GovernedStateHeadPreconditionV1Alpha1, + ImmutableRecordReferenceV1, + OutcomeIntentV1Alpha1, + OutcomeV1Alpha1, + canonical_hash, + canonical_json, + capability_state_ref_for_artifact, +) +from ace.intelligence import ( + CanonicalJsonValueV1Alpha1, + EvidenceAcquisitionMode, + ImpactClassification, + ImpactConditionsV1Alpha1, + ImpactCriterionV1Alpha1, + ImpactEvaluationRequestV1Alpha1, + ImpactEvidenceV1Alpha1, + ImpactGovernanceAction, + ImpactMetricDirection, + ImpactOutcomeMeasuresV1Alpha1, + ImpactTargetKind, + IntelligenceResourceMode, + ObservationV1Alpha1, +) +from pydantic import Field, field_validator, model_validator + +from scripts.p2c2_federal_register_monitor import _time +from scripts.p2c2_governed_reality_brief import _context, _head +from scripts.p2c3_measured_feedback import ( + ReviewedExport, + _append_value, + _authorize_append, + _digest, + _record_attribution, + _run_reviewed_export, +) +from scripts.p2c5_citation_correctness_outcome import _derive_identity, _FrozenModel +from scripts.p2c7_correction_detection_delay_outcome import _load_observation +from scripts.p2c9_forecast_calibration_outcome import run_forecast_calibration_outcome + +FIXTURE_PATH = ( + Path(__file__).resolve().parents[1] / "domain_packs" / "tests" / "fixtures" / "p2c10_bls_correction_pair.json" +) +OUTCOME_TYPE = "independent_artifact_review" +MEASURE_ID = "official_correction_statement_quality" +CRITERION_ID = "impact_criterion:world-independent-official-correction-statement-quality" +CRITERION_FROZEN_AT = _time("2026-08-10T23:59:00Z") +REVIEW_POLICY_ID = "world_independent_official_correction_statement_quality" +REVIEW_POLICY_VERSION = "candidate-1" + +ORIGINAL_STATEMENT = "The number of job openings decreased in federal government (39,000)." +CORRECTED_STATEMENT = "The number of job openings decreased in federal government (−39,000)." +RELEASE_URI = "https://www.bls.gov/news.release/archives/jolts_07012025.htm" +ERRATA_URI = "https://www.bls.gov/errata/" + +IMPACT_ARTIFACT = CapabilityArtifactIdentityV1Alpha1( + capability="measured_impact_evaluation", + contract="ace.application.measured-impact-service/v1alpha1", + implementation_id="world_independent_correction_reproduction_candidate", + implementation_version="0.1.0", + artifact_digest="sha256:" + "d" * 64, +) + + +class OfficialCorrectionArtifactV1Alpha1(_FrozenModel): + """One exact World rendering of an official correction pair.""" + + contract: Literal["ace.world-intelligence.official-correction-artifact/v1alpha1"] = ( + "ace.world-intelligence.official-correction-artifact/v1alpha1" + ) + product_id: str + artifact_key: str + original_observation: ImmutableRecordReferenceV1 + correction_observation: ImmutableRecordReferenceV1 + source_family: Literal["bls_public_errata"] = "bls_public_errata" + source_release_id: str + correction_record_id: str + corrects_source_record_id: str + correction_relation: Literal["corrects"] = "corrects" + original_statement: str + corrected_statement: str + displayed_statement: str + source_coverage_complete: Literal[True] = True + correction_link_visible: Literal[True] = True + prior_record_preserved: Literal[True] = True + limitations: tuple[str, ...] + generated_at: datetime + artifact_id: str | None = None + artifact_digest: str | None = None + + @field_validator("limitations") + @classmethod + def canonicalize_limitations(cls, value: tuple[str, ...]) -> tuple[str, ...]: + ordered = tuple(sorted(value)) + if not ordered or len(ordered) != len(set(ordered)): + raise ValueError("official-correction limitations must be non-empty and unique") + return ordered + + @model_validator(mode="after") + def validate_scope_pair_and_identity(self) -> Self: + if ( + self.original_observation.product_id != self.product_id + or self.correction_observation.product_id != self.product_id + ): + raise ValueError("official-correction artifact crossed exact product scope") + if self.original_observation.storage_id == self.correction_observation.storage_id: + raise ValueError("official-correction artifact requires distinct exact source records") + if self.corrects_source_record_id != self.source_release_id: + raise ValueError("official-correction artifact lost its exact correction linkage") + if self.original_statement != ORIGINAL_STATEMENT or self.corrected_statement != CORRECTED_STATEMENT: + raise ValueError("official-correction artifact changed the frozen statement pair") + if self.displayed_statement not in {self.original_statement, self.corrected_statement}: + raise ValueError("official-correction artifact introduced an unreviewed statement form") + _derive_identity( + self, + prefix="official_correction_artifact", + id_field="artifact_id", + digest_field="artifact_digest", + ) + return self + + +class IndependentCorrectionReviewV1Alpha1(_FrozenModel): + """Exact product-owned review of correction visibility and statement quality.""" + + contract: Literal["ace.world-intelligence.independent-correction-review/v1alpha1"] = ( + "ace.world-intelligence.independent-correction-review/v1alpha1" + ) + product_id: str + review_key: str + reviewed_subject: ImmutableRecordReferenceV1 + original_observation: ImmutableRecordReferenceV1 + correction_observation: ImmutableRecordReferenceV1 + reviewer_context: AuthenticatedRuntimeContextV1Alpha1 + policy_id: str + policy_version: str + policy_digest: str + source_fixture_digest: str + source_family: Literal["bls_public_errata"] = "bls_public_errata" + source_coverage_complete: bool + correction_link_visible: bool + prior_record_preserved: bool + original_statement: str + expected_corrected_statement: str + displayed_statement: str + corrected_statement_exact: bool + stale_form_present: bool + correction_quality_score: float = Field(ge=0.0, le=1.0) + limitations: tuple[str, ...] + reviewed_at: datetime + review_id: str | None = None + review_digest: str | None = None + + @field_validator("limitations") + @classmethod + def canonicalize_limitations(cls, value: tuple[str, ...]) -> tuple[str, ...]: + ordered = tuple(sorted(value)) + if not ordered or len(ordered) != len(set(ordered)): + raise ValueError("independent-correction review limitations must be non-empty and unique") + return ordered + + @model_validator(mode="after") + def validate_scope_score_and_identity(self) -> Self: + references = (self.reviewed_subject, self.original_observation, self.correction_observation) + if any(item.product_id != self.product_id for item in references): + raise ValueError("independent-correction review crossed exact product scope") + if self.reviewer_context.product_id != self.product_id: + raise ValueError("independent-correction reviewer crossed exact product scope") + if self.original_observation.storage_id == self.correction_observation.storage_id: + raise ValueError("independent-correction review requires distinct source records") + exact = self.displayed_statement == self.expected_corrected_statement + stale = self.displayed_statement == self.original_statement + if self.corrected_statement_exact != exact or self.stale_form_present != stale: + raise ValueError("independent-correction statement disposition was not derived exactly") + expected_score = float( + self.source_coverage_complete + and self.correction_link_visible + and self.prior_record_preserved + and exact + and not stale + ) + if self.correction_quality_score != expected_score: + raise ValueError("correction quality score differs from the frozen product rule") + _derive_identity( + self, + prefix="independent_correction_review", + id_field="review_id", + digest_field="review_digest", + ) + return self + + +class _ReplayMustNotAuthorize: + async def authorize_action(self, request): + raise AssertionError( + f"historical independent-correction evaluation requested new authority: {request.authorization_key}" + ) + + +def validate_bls_correction_fixture(fixture: dict[str, Any]) -> dict[str, Any]: + """Fail closed if the recorded BLS source pair or its policy drifts.""" + + if fixture.get("network_access") is not False: + raise AssertionError("recorded BLS correction fixture must remain network-free") + if fixture.get("source_policy") != { + "publisher": "U.S. Bureau of Labor Statistics", + "source_family": "bls_public_errata", + "release_uri": RELEASE_URI, + "errata_uri": ERRATA_URI, + "recorded_replay": True, + "historical_original_form_derived_from_erratum": True, + "statistical_validity_claimed": False, + }: + raise AssertionError("recorded BLS correction source policy changed") + original = fixture.get("original", {}) + correction = fixture.get("correction", {}) + if original.get("record_id") != "USDL-25-1087" or original.get("release_uri") != RELEASE_URI: + raise AssertionError("recorded BLS release identity changed") + if original.get("reported_sentence_without_required_minus_sign") != ORIGINAL_STATEMENT: + raise AssertionError("recorded BLS original statement changed") + if ( + correction.get("record_id") != "bls-errata-2025-07-01-jolts" + or correction.get("corrects_record_id") != original.get("record_id") + or correction.get("errata_uri") != ERRATA_URI + ): + raise AssertionError("recorded BLS correction linkage changed") + if correction.get("corrected_sentence") != CORRECTED_STATEMENT: + raise AssertionError("recorded BLS corrected statement changed") + return fixture + + +def load_bls_correction_fixture() -> dict[str, Any]: + return validate_bls_correction_fixture(json.loads(FIXTURE_PATH.read_text(encoding="utf-8"))) + + +def bls_correction_fixture_digest(fixture: dict[str, Any]) -> str: + return f"sha256:{canonical_hash(fixture)}" + + +def _install_policy(state: dict[str, Any]): + environment = state["environment"] + runtime = state["runtime"] + product_id = environment.fixture["product_id"] + criterion_head = _head(product_id, "impact_criterion", CRITERION_ID, 110) + operation_head = _head( + product_id, + "governed_operation_configuration", + "governed_operation_configuration:world-independent-official-correction-quality", + 111, + ) + binding = GovernedOperationBindingV1Alpha1( + product_id=product_id, + artifact=IMPACT_ARTIFACT, + configuration_ref=operation_head.state_id, + authority="append_measured_impact", + grant_ref="authority_grant:world-independent-official-correction-quality", + state_head_precondition=GovernedStateHeadPreconditionV1Alpha1.from_head(operation_head), + ) + capability_head = _head( + product_id, + "capability_state", + capability_state_ref_for_artifact(IMPACT_ARTIFACT), + 112, + ) + authority_head = _head(product_id, "authority_grant", binding.grant_ref, 113) + for head in (criterion_head, operation_head, capability_head, authority_head): + environment.store.set_governed_state_head(head) + runtime.heads[head.state_kind, head.state_id] = head + runtime.bindings = (*runtime.bindings, binding) + return criterion_head, binding + + +async def _append_source_observation( + state: dict[str, Any], + *, + fixture: dict[str, Any], + role: Literal["original", "correction"], +) -> tuple[ObservationV1Alpha1, ImmutableRecordReferenceV1]: + environment = state["environment"] + source = fixture[role] + source_id = source["record_id"] + published_at = _time(source["published_at"] if role == "original" else source["corrected_at"]) + ingested_at = _time(fixture["recorded_at"]) + fixture_digest = bls_correction_fixture_digest(fixture) + payload = { + "fixture_id": fixture["fixture_id"], + "fixture_digest": fixture_digest, + "record_role": role, + "source_policy": fixture["source_policy"], + "record": source, + } + observation = ObservationV1Alpha1( + product_id=environment.fixture["product_id"], + mode=IntelligenceResourceMode.PREPARED, + activation_revision=state["brief_admission"].brief.activation_revision, + as_of=ingested_at, + source_ref=f"bls_{'release' if role == 'original' else 'erratum'}:{source_id}", + source_digest=_digest(source), + acquisition_mode=EvidenceAcquisitionMode.RECORDED_REPLAY, + acquisition_receipt_ref=f"recorded_replay_acquisition:{source_id}", + acquisition_receipt_digest=_digest( + {"fixture_digest": fixture_digest, "record_id": source_id, "network_access": False} + ), + source_published_at=published_at, + event_effective_at=None, + observed_at=published_at, + ingested_at=ingested_at, + subject_refs=("bls_program:jolts", "bls_release:USDL-25-1087"), + payload=CanonicalJsonValueV1Alpha1(value_json=canonical_json(payload)), + confidence=1.0, + ) + requested_at = state["clock"]() + authorization = await _authorize_append( + state, + context=environment.context, + authorization_key=f"append:world-bls-correction-observation:{role}", + subject_ref=str(observation.resource_id), + subject_digest=str(observation.resource_digest), + requested_at=requested_at, + ) + reference = await _append_value( + state, + value=observation, + record_kind="observation", + record_key=str(observation.resource_id), + transaction_key=f"world-bls-correction-observation:{observation.resource_id}", + as_of=observation.as_of, + authorization=authorization, + ) + return observation, reference + + +async def _append_artifact( + state: dict[str, Any], + *, + fixture: dict[str, Any], + original_ref: ImmutableRecordReferenceV1, + correction_ref: ImmutableRecordReferenceV1, + variant: Literal["treatment", "control"], +) -> tuple[OfficialCorrectionArtifactV1Alpha1, ImmutableRecordReferenceV1, str]: + environment = state["environment"] + displayed = CORRECTED_STATEMENT if variant == "treatment" else ORIGINAL_STATEMENT + artifact = OfficialCorrectionArtifactV1Alpha1( + product_id=environment.fixture["product_id"], + artifact_key=f"bls-jolts-correction:{variant}:USDL-25-1087", + original_observation=original_ref, + correction_observation=correction_ref, + source_release_id=fixture["original"]["record_id"], + correction_record_id=fixture["correction"]["record_id"], + corrects_source_record_id=fixture["correction"]["corrects_record_id"], + original_statement=ORIGINAL_STATEMENT, + corrected_statement=CORRECTED_STATEMENT, + displayed_statement=displayed, + limitations=( + "historical_original_form_derived_from_public_erratum", + "one_recorded_bls_correction_not_population_performance", + "recorded_replay_not_live_monitoring", + "statement_quality_not_statistical_validity_or_human_benefit", + ), + generated_at=state["clock"](), + ) + authorization = await _authorize_append( + state, + context=environment.context, + authorization_key=f"append:world-bls-correction-artifact:{variant}", + subject_ref=str(artifact.artifact_id), + subject_digest=str(artifact.artifact_digest), + requested_at=state["clock"](), + ) + reference = await _append_value( + state, + value=artifact, + record_kind="official_correction_artifact", + record_key=str(artifact.artifact_id), + transaction_key=f"world-bls-correction-artifact:{artifact.artifact_id}", + as_of=artifact.generated_at, + authorization=authorization, + ) + content = ( + "# BLS JOLTS Correction Review\n\n" + f"Original Observation: {original_ref.storage_id}\n\n" + f"Correction Observation: {correction_ref.storage_id}\n\n" + f"Release: {RELEASE_URI}\n\n" + f"Erratum: {ERRATA_URI}\n\n" + f"Reviewed statement: {displayed}\n" + ) + return artifact, reference, content + + +async def _load_artifact( + state: dict[str, Any], reference: ImmutableRecordReferenceV1 +) -> OfficialCorrectionArtifactV1Alpha1: + record = await state["environment"].store.load_record( + reference.storage_id, + product_id=reference.product_id, + record_space=reference.record_space, + record_kind=reference.record_kind, + ) + if ( + record is None + or record.reference() != reference + or record.payload_contract != "ace.world-intelligence.official-correction-artifact/v1alpha1" + ): + raise AssertionError("official-correction artifact is unavailable or changed") + return OfficialCorrectionArtifactV1Alpha1.model_validate(record.payload) + + +def _policy_digest( + fixture: dict[str, Any], + *, + original_ref: ImmutableRecordReferenceV1, + correction_ref: ImmutableRecordReferenceV1, +) -> str: + return _digest( + { + "policy_id": REVIEW_POLICY_ID, + "policy_version": REVIEW_POLICY_VERSION, + "fixture_digest": bls_correction_fixture_digest(fixture), + "original_observation": original_ref.model_dump(mode="json"), + "correction_observation": correction_ref.model_dump(mode="json"), + "original_statement": ORIGINAL_STATEMENT, + "corrected_statement": CORRECTED_STATEMENT, + "score": ( + "1 only when exact source coverage, correction linkage, prior-record preservation, " + "corrected statement, and stale-form removal all pass; otherwise 0" + ), + } + ) + + +async def _review_artifact( + state: dict[str, Any], + *, + fixture: dict[str, Any], + subject_ref: ImmutableRecordReferenceV1, + original_ref: ImmutableRecordReferenceV1, + correction_ref: ImmutableRecordReferenceV1, + pair_index: int, + variant: Literal["treatment", "control"], +) -> tuple[IndependentCorrectionReviewV1Alpha1, ImmutableRecordReferenceV1]: + environment = state["environment"] + artifact = await _load_artifact(state, subject_ref) + original = await _load_observation(state, original_ref) + correction = await _load_observation(state, correction_ref) + original_payload = original.payload.parsed_value() + correction_payload = correction.payload.parsed_value() + expected_original = original_payload["record"] == fixture["original"] + expected_correction = correction_payload["record"] == fixture["correction"] + source_coverage = bool( + expected_original + and expected_correction + and artifact.original_observation == original_ref + and artifact.correction_observation == correction_ref + ) + linkage = bool( + artifact.correction_relation == "corrects" + and artifact.corrects_source_record_id == artifact.source_release_id + and fixture["correction"]["corrects_record_id"] == fixture["original"]["record_id"] + ) + prior_preserved = bool(await _load_observation(state, original_ref) == original) + exact = artifact.displayed_statement == CORRECTED_STATEMENT + stale = artifact.displayed_statement == ORIGINAL_STATEMENT + reviewer = _context(environment.context, "principal:world-independent-correction-reviewer") + review = IndependentCorrectionReviewV1Alpha1( + product_id=environment.fixture["product_id"], + review_key=f"bls-independent-correction-quality:{pair_index}:{variant}", + reviewed_subject=subject_ref, + original_observation=original_ref, + correction_observation=correction_ref, + reviewer_context=reviewer, + policy_id=REVIEW_POLICY_ID, + policy_version=REVIEW_POLICY_VERSION, + policy_digest=_policy_digest( + fixture, + original_ref=original_ref, + correction_ref=correction_ref, + ), + source_fixture_digest=bls_correction_fixture_digest(fixture), + source_coverage_complete=source_coverage, + correction_link_visible=linkage, + prior_record_preserved=prior_preserved, + original_statement=ORIGINAL_STATEMENT, + expected_corrected_statement=CORRECTED_STATEMENT, + displayed_statement=artifact.displayed_statement, + corrected_statement_exact=exact, + stale_form_present=stale, + correction_quality_score=float(source_coverage and linkage and prior_preserved and exact and not stale), + limitations=artifact.limitations, + reviewed_at=state["clock"](), + ) + authorization = await _authorize_append( + state, + context=reviewer, + authorization_key=f"review:world-bls-independent-correction:{pair_index}:{variant}", + subject_ref=str(review.review_id), + subject_digest=str(review.review_digest), + requested_at=state["clock"](), + ) + reference = await _append_value( + state, + value=review, + record_kind="independent_correction_review", + record_key=str(review.review_id), + transaction_key=f"world-independent-correction-review:{review.review_id}", + as_of=review.reviewed_at, + authorization=authorization, + ) + return review, reference + + +async def _record_review_outcome( + state: dict[str, Any], + *, + export: ReviewedExport, + review: IndependentCorrectionReviewV1Alpha1, + review_ref: ImmutableRecordReferenceV1, + pair_index: int, + variant: Literal["treatment", "control"], +) -> ImmutableRecordReferenceV1: + environment = state["environment"] + latency_ms = max(0, int((review.reviewed_at - export.intent.requested_at).total_seconds() * 1_000)) + measures = ImpactOutcomeMeasuresV1Alpha1( + primary_value=review.correction_quality_score, + observed_result=review_ref, + latency_ms=latency_ms, + cost_usd=0.0, + failure_count=0, + degraded=False, + limitations=review.limitations, + ) + observer = _context(environment.context, "principal:world-independent-correction-outcome-observer") + recorded_at = state["clock"]() + intent = OutcomeIntentV1Alpha1( + product_id=environment.fixture["product_id"], + authenticated_context=observer, + decision=export.decision_ref, + outcome_type=OUTCOME_TYPE, + measure_id=MEASURE_ID, + value_json=canonical_json(measures.model_dump(mode="json")), + observed_at=review.reviewed_at, + recorded_at=recorded_at, + ) + authorization = await _authorize_append( + state, + context=observer, + authorization_key=f"outcome:world-bls-independent-correction:{pair_index}:{variant}", + subject_ref=str(intent.intent_id), + subject_digest=str(intent.intent_digest), + requested_at=recorded_at, + ) + outcome = OutcomeV1Alpha1(intent=intent, authorization=authorization) + return await _append_value( + state, + value=outcome, + record_kind="outcome", + record_key=str(outcome.outcome_id), + transaction_key=f"outcome:{outcome.outcome_id}", + as_of=review.reviewed_at, + authorization=authorization, + ) + + +async def run_independent_correction_reproduction(workspace_root: Path) -> dict[str, Any]: + """Run P2C10 over a recorded BLS correction and unchanged impact contracts.""" + + state: dict[str, Any] = {} + prior_packet = await run_forecast_calibration_outcome(workspace_root, state_sink=state) + environment = state["environment"] + fixture = load_bls_correction_fixture() + original, original_ref = await _append_source_observation(state, fixture=fixture, role="original") + correction, correction_ref = await _append_source_observation(state, fixture=fixture, role="correction") + criterion_head, impact_binding = _install_policy(state) + treatment_artifact, treatment_ref, treatment_content = await _append_artifact( + state, + fixture=fixture, + original_ref=original_ref, + correction_ref=correction_ref, + variant="treatment", + ) + control_artifact, control_ref, control_content = await _append_artifact( + state, + fixture=fixture, + original_ref=original_ref, + correction_ref=correction_ref, + variant="control", + ) + treatment_exports = tuple( + [ + await _run_reviewed_export( + state, + subject=treatment_ref, + content=treatment_content, + pair_index=index, + variant="bls-independent-correction-treatment", + ) + for index in (1, 2) + ] + ) + control_exports = tuple( + [ + await _run_reviewed_export( + state, + subject=control_ref, + content=control_content, + pair_index=index, + variant="bls-independent-correction-control", + ) + for index in (1, 2) + ] + ) + + evidence: list[ImpactEvidenceV1Alpha1] = [] + treatment_reviews: list[IndependentCorrectionReviewV1Alpha1] = [] + control_reviews: list[IndependentCorrectionReviewV1Alpha1] = [] + for index, (treatment_export, control_export) in enumerate( + zip(treatment_exports, control_exports, strict=True), start=1 + ): + treatment_attribution = await _record_attribution( + state, + export=treatment_export, + subject=treatment_ref, + pair_index=index, + variant="bls-independent-correction-treatment", + ) + control_attribution = await _record_attribution( + state, + export=control_export, + subject=control_ref, + pair_index=index, + variant="bls-independent-correction-control", + ) + treatment_review, treatment_review_ref = await _review_artifact( + state, + fixture=fixture, + subject_ref=treatment_ref, + original_ref=original_ref, + correction_ref=correction_ref, + pair_index=index, + variant="treatment", + ) + control_review, control_review_ref = await _review_artifact( + state, + fixture=fixture, + subject_ref=control_ref, + original_ref=original_ref, + correction_ref=correction_ref, + pair_index=index, + variant="control", + ) + treatment_outcome = await _record_review_outcome( + state, + export=treatment_export, + review=treatment_review, + review_ref=treatment_review_ref, + pair_index=index, + variant="treatment", + ) + control_outcome = await _record_review_outcome( + state, + export=control_export, + review=control_review, + review_ref=control_review_ref, + pair_index=index, + variant="control", + ) + treatment_reviews.append(treatment_review) + control_reviews.append(control_review) + conditions = ImpactConditionsV1Alpha1( + product_id=environment.fixture["product_id"], + condition_key=f"world-bls-independent-correction-quality-pair:{index}", + route_id="world:bls-recorded-correction-review", + context_json=canonical_json( + { + "fixture_digest": bls_correction_fixture_digest(fixture), + "pair_index": index, + "review_policy_digest": treatment_review.policy_digest, + "source_family": "bls_public_errata", + "source_refs": sorted((original_ref.storage_id, correction_ref.storage_id)), + "task": "render_exact_official_correction_with_prior_record_preserved", + } + ), + observation_window_start=CRITERION_FROZEN_AT, + observation_window_end=state["clock"](), + frozen_at=CRITERION_FROZEN_AT, + ) + evidence.append( + ImpactEvidenceV1Alpha1( + product_id=environment.fixture["product_id"], + evidence_key=f"world-bls-independent-correction-quality:{index}", + treatment_attribution=treatment_attribution, + control_attribution=control_attribution, + treatment_decision=treatment_export.decision_ref, + control_decision=control_export.decision_ref, + treatment_action_review=treatment_export.review_ref, + treatment_action_admission=treatment_export.admission_ref, + treatment_action_terminal=treatment_export.terminal_ref, + control_action_review=control_export.review_ref, + control_action_admission=control_export.admission_ref, + control_action_terminal=control_export.terminal_ref, + treatment_outcome=treatment_outcome, + control_outcome=control_outcome, + treatment_conditions=conditions, + control_conditions=conditions, + ) + ) + + criterion = ImpactCriterionV1Alpha1( + product_id=environment.fixture["product_id"], + criterion_id=CRITERION_ID, + criterion_version="candidate-1", + target_kind=ImpactTargetKind.INTELLIGENCE_ARTIFACT, + outcome_type=OUTCOME_TYPE, + measure_id=MEASURE_ID, + metric_direction=ImpactMetricDirection.HIGHER_IS_BETTER, + useful_effect_threshold=1.0, + harmful_effect_threshold=1.0, + minimum_matched_pairs=2, + requires_observed_result=True, + harmful_action=ImpactGovernanceAction.ROLLBACK, + state_head_precondition=GovernedStateHeadPreconditionV1Alpha1.from_head(criterion_head), + frozen_at=CRITERION_FROZEN_AT, + ) + request = ImpactEvaluationRequestV1Alpha1( + evaluation_key="world-independent-correction-quality:bls-jolts-2025-minus-sign", + product_id=environment.fixture["product_id"], + authenticated_context=environment.context, + criterion=criterion, + target=treatment_ref, + control=control_ref, + evidence=tuple(evidence), + cutoff_at=state["clock"](), + requested_at=state["clock"](), + ) + service = MeasuredImpactService( + store=environment.store, + authorizer=state["reasoning"], + operation_binding=impact_binding, + ) + admission = await service.evaluate(request) + historical = await MeasuredImpactService( + store=environment.store, + authorizer=_ReplayMustNotAuthorize(), + operation_binding=impact_binding, + ).evaluate(request) + if admission.replayed or not historical.replayed or historical.evaluation != admission.evaluation: + raise AssertionError("independent-correction evaluation did not replay exact historical material") + if admission.evaluation.classification is not ImpactClassification.USEFUL: + raise AssertionError("frozen independent-correction criterion did not classify useful") + if admission.proposal is None or admission.proposal.action is not ImpactGovernanceAction.PROMOTE: + raise AssertionError("independent-correction result did not emit its proposal-only mapping") + if {item.correction_quality_score for item in treatment_reviews} != {1.0}: + raise AssertionError("independent-correction treatment lost the exact corrected statement") + if {item.correction_quality_score for item in control_reviews} != {0.0}: + raise AssertionError("independent-correction control did not retain the stale statement form") + + state.update( + { + "p2c10_fixture": fixture, + "p2c10_original_observation": original, + "p2c10_original_observation_ref": original_ref, + "p2c10_correction_observation": correction, + "p2c10_correction_observation_ref": correction_ref, + } + ) + return { + "contract": "ace.world-intelligence.independent-correction-reproduction/v1alpha1", + "prior_forecast_calibration": prior_packet, + "source_pair": { + "fixture_id": fixture["fixture_id"], + "fixture_digest": bls_correction_fixture_digest(fixture), + "network_access": fixture["network_access"], + "source_policy": fixture["source_policy"], + "original": fixture["original"], + "correction": fixture["correction"], + "original_observation": original_ref.model_dump(mode="json"), + "correction_observation": correction_ref.model_dump(mode="json"), + }, + "review_policy": { + "policy_id": REVIEW_POLICY_ID, + "policy_version": REVIEW_POLICY_VERSION, + "policy_digest": treatment_reviews[0].policy_digest, + "reviewer_ref": treatment_reviews[0].reviewer_context.actor_ref, + "score_rule": ( + "1 only when exact source coverage, correction linkage, prior-record preservation, " + "corrected statement, and stale-form removal all pass; otherwise 0" + ), + }, + "artifacts": { + "treatment": treatment_artifact.model_dump(mode="json"), + "control": control_artifact.model_dump(mode="json"), + }, + "observed_results": { + "treatment": tuple(item.model_dump(mode="json") for item in treatment_reviews), + "control": tuple(item.model_dump(mode="json") for item in control_reviews), + }, + "evaluation": admission.evaluation.model_dump(mode="json"), + "proposal": admission.proposal.model_dump(mode="json"), + "replay": { + "historical": historical.replayed, + "no_reauthorization": True, + "transaction_receipt_id": str(historical.transaction_receipt.receipt_id), + }, + "scope": { + "independent_source_family_reproduction": True, + "exact_recorded_official_correction_pair": True, + "prior_record_preserved": True, + "domain_neutral_core_contract_unchanged": True, + "domain_pack_changed": False, + "connector_changed": False, + "network_access": False, + "live_monitoring_claimed": False, + "statistical_validity_claimed": False, + "population_correction_performance_claimed": False, + "causality_claimed": False, + "human_benefit_claimed": False, + "proposal_applied": False, + "autonomous_publication": False, + }, + } + + +def main() -> None: + import argparse + + parser = argparse.ArgumentParser() + parser.add_argument("workspace_root", type=Path) + args = parser.parse_args() + print( + json.dumps( + asyncio.run(run_independent_correction_reproduction(args.workspace_root)), + indent=2, + sort_keys=True, + ) + ) + + +if __name__ == "__main__": + main() diff --git a/scripts/p2c9_forecast_calibration_outcome.py b/scripts/p2c9_forecast_calibration_outcome.py index 047afab..cb82693 100644 --- a/scripts/p2c9_forecast_calibration_outcome.py +++ b/scripts/p2c9_forecast_calibration_outcome.py @@ -551,10 +551,14 @@ async def _record_review_outcome( ) -async def run_forecast_calibration_outcome(workspace_root: Path) -> dict[str, Any]: +async def run_forecast_calibration_outcome( + workspace_root: Path, + *, + state_sink: dict[str, Any] | None = None, +) -> dict[str, Any]: """Run P2C9 over a withheld exact correction result and declared probabilities.""" - state: dict[str, Any] = {} + state: dict[str, Any] = {} if state_sink is None else state_sink prior_packet = await run_correction_revision_stability_outcome( workspace_root, state_sink=state,