diff --git a/README.md b/README.md index e86e10b..d40a6c3 100644 --- a/README.md +++ b/README.md @@ -238,6 +238,16 @@ The recorded availability and detection instants are test coordinates. The suite network access and does not establish live monitoring, network-arrival latency, population delay performance, legal truth, calibration, general Brief quality, causality, or human benefit. +P2C8 creates actual prior, treatment, and control `BriefV1Alpha1` records over that same source +pair. Both revisions remove the stale instruction claim, add the same exact correction claim, +retain three claims, and cite the same original/correction sources. Treatment preserves the exact +identities of both unaffected claims; the paraphrase-drift control preserves neither. Independent +reviews record every expected, preserved, drifted, and unexpected claim identity, producing scores +of `1.0` and `0.0` across two matched pairs and another non-effective proposal. + +This freezes a World product rule over one recorded pair. It is not live revision, a general +semantic-equivalence engine, population stability, calibration, causality, or human benefit. + ## What the public World proof demonstrates Generate a self-contained visual Reality Brief and its exact machine-readable backing data: @@ -332,10 +342,11 @@ $PY -m scripts.p2c3_measured_feedback "$WORKSPACE" # Stacked candidate: explicit reject/no-action review of the exact proposal $PY -m scripts.p2c4_reviewed_impact_disposition "$WORKSPACE" -# Stacked candidates: exact citation correctness, contradiction attention, and correction delay +# Stacked candidates: correctness, attention, correction delay, and revision stability $PY -m scripts.p2c5_citation_correctness_outcome "$WORKSPACE" $PY -m scripts.p2c6_contradiction_attention_outcome "$WORKSPACE" $PY -m scripts.p2c7_correction_detection_delay_outcome "$WORKSPACE" +$PY -m scripts.p2c8_correction_revision_stability_outcome "$WORKSPACE" ``` The released 0.9.0 gates are reproducible through the locked environment, as CI does. The @@ -404,9 +415,10 @@ adds an independently reviewed citation-correctness Outcome and a citation-prese negative control. P2C6 adds exact contradiction recall, false-alert rate, equal alert-volume control, and valid silence under one frozen recorded-source challenge. P2C7 adds exact handling of one explicit recorded correction pair, preserves the prior record, and measures frozen-replay -detection delay against a product target. The next bounded measurement work is -calibration/revision stability, another independently sourced correction event, or independent -Market reproduction. +detection delay against a product target. P2C8 measures exact unaffected-claim identity +preservation across real Brief contracts while holding correction semantics, source coverage, and +claim count constant. The next bounded measurement work is calibration, another independently +sourced correction event, or independent Market reproduction. Separately reviewed opt-in network transport and P2D multi-source conflict/correction with LIVE inputs remain independent work. None of these steps may add autonomous publishing, delivery, persuasion, or action authority to a Domain Pack. diff --git a/ROADMAP.md b/ROADMAP.md index 59c5f92..70ca3f4 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -16,7 +16,7 @@ source code into the platform. See the [0.9.0 release record](docs/releases/world-intelligence-p2c2-v0.9.0.md). -## Candidate — P2C3–P2C7 measured feedback and product-owned outcomes +## Candidate — P2C3–P2C8 measured feedback and product-owned outcomes - The exact 0.9.0 Brief and a World-owned source-only control pass through two matched reviewed export pairs under one frozen structural citation-coverage criterion. @@ -41,6 +41,11 @@ See the [0.9.0 release record](docs/releases/world-intelligence-p2c2-v0.9.0.md). reviewed workflow, but their replay delays are 300 and 21600 seconds against a product-owned 600-second target. Exact reviews and Outcomes expose the source pair, rule, delay, score, and limitations without claiming live network-arrival performance. +- A fifth frozen criterion creates actual prior and revised Brief contracts over that correction. + Treatment and control have the same exact sources, replacement claim, claim count, and reviewed + workflow, but treatment preserves both unaffected claim identities while the paraphrase-drift + control preserves neither. Exact reviews expose the affected, replacement, stable, preserved, + drifted, and unexpected claim sets without claiming general semantic equivalence. - This is source-checkout evidence against stacked Core candidates, not a released World capability, human benefit finding, causal claim, network-freshness proof, or applied governance change. @@ -55,7 +60,10 @@ freezes contradiction recall, false-alert rate, equal alert volume, and valid si claiming live conflict detection or population performance. The stacked [P2C7 work packet](docs/design/world-intelligence-p2c7-correction-detection-delay-outcome-work-packet-v1.md) freezes exact correction linkage, prior-record preservation, and recorded-replay detection delay -without claiming live monitoring or network-arrival latency. +without claiming live monitoring or network-arrival latency. The stacked +[P2C8 work packet](docs/design/world-intelligence-p2c8-correction-revision-stability-outcome-work-packet-v1.md) +freezes correction-induced Brief revision stability without claiming live revision, population +performance, or a general semantic-equivalence engine. ## Next — trustworthy live orientation diff --git a/docs/audits/world-intelligence-p2c8-correction-revision-stability-outcome-2026-08-10.md b/docs/audits/world-intelligence-p2c8-correction-revision-stability-outcome-2026-08-10.md new file mode 100644 index 0000000..22ca9ce --- /dev/null +++ b/docs/audits/world-intelligence-p2c8-correction-revision-stability-outcome-2026-08-10.md @@ -0,0 +1,154 @@ +# World Intelligence P2C8 correction-revision-stability outcome audit — 2026-08-10 + +Status: **stacked candidate evidence only; not a release, SI4 pass, live-revision proof, or applied +governance change** + +## Source identity + +- World base: P2C7 commit `216a37a1fdcb4f0baf9ac148ab9e525559141c22` +- World branch: `codex/world-correction-revision-stability` +- Core dependency: exact observed-result candidate commit + `433e3d16c5458c975557dcd1552824fb959d4d12` +- Released identity intentionally unchanged: `ace-domain-world-intelligence==0.9.0`, + `ace-core>=0.5.0,<0.6` + +## Exact point-in-time result + +The source fixture remains +`sha256:2b81d3950cbfd127408eec227ec5cd249677a189120d6ca7b603d85d01074543`, +with original Observation `observation:f1768d6f4191a86e245846a9a1e33768` and correction +Observation `observation:fced5d3bbc3802c0285021142b332e29`. + +One source-checkout run recorded actual prior, treatment, and control `BriefV1Alpha1` resources: + +```text +prior Brief: + brief:b0911c340cc4a39cd7a908f7884d71bd + sha256:b0911c340cc4a39cd7a908f7884d71bd6348b6202b9d8f521d4339e708273250 +treatment Brief: + brief:d1dafb5352ebfd99bc35c1044bc627ac + sha256:d1dafb5352ebfd99bc35c1044bc627ac3e30846b8092561ec28bfca4ae4921c3 +control Brief: + brief:f86f101516ea496108557498e5950672 + sha256:f86f101516ea496108557498e5950672660ead2c1a01949a46776250cbf3f075 + +affected prior claim: + grounded_claim:fddf5435d53f01312739f8f8ff355eb6 +replacement correction claim: + grounded_claim:ff1683b3155f734bcf942c0cab9ed7e4 +stable claims: + grounded_claim:7c6c683cae2dabeb050d5542405e992a + grounded_claim:8fab51672b58cdd0c45734ea637d7aaf +control-only paraphrase identities: + grounded_claim:82b05b5ec627f03a43e76c76eff27fde + grounded_claim:ec5fc4e83b25519dc44f2878ed69588b + +original citation: citation:0cb9a00ce581b4d09a0ab14f755caf05 +correction citation: citation:13ae2675d1ce8f512389eda12e4b6632 +``` + +Treatment and control both have three claims, both remove the affected prior claim, both add the +same replacement, and both cite the same original and correction sources. Treatment preserves both +stable claim identities. Control preserves neither and introduces the two frozen paraphrase +identities. + +The reviews used product policy `world_recorded_correction_revision_stability` version +`candidate-1`, material +`sha256:0145cc35244c21b84f3ba346ccb16192b825fd91d1bcff950e3f0a8005ac932b`: + +```text +treatment review 1: brief_revision_stability_review:3b50050e155af3d3eb4391e232a178d0 +treatment review 2: brief_revision_stability_review:9f0544d7c79b9a6601760d98133db156 +control review 1: brief_revision_stability_review:ac1422dde272aed347045d671e2c6276 +control review 2: brief_revision_stability_review:1a4d0d821b3752a47d15a3d6a1746bc7 +treatment preserved stable claims: 2, 2 +control preserved stable claims: 0, 0 +treatment drifted stable claims: 0, 0 +control drifted stable claims: 2, 2 +treatment score: 1.0, 1.0 +control score: 0.0, 0.0 +matched pairs: 2 +mean effect: 1.0 +classification: useful +proposal action: promote +proposal live effect: false +historical replay: true +replay reauthorization: false +``` + +The exact evaluation was `impact_evaluation:d0ba0242b75dd5ec63ba27aa223a0225` with material +`sha256:d0ba0242b75dd5ec63ba27aa223a0225ff83239a40398a8cbd50d54cea1a3a0a`. +The exact non-effective proposal was +`impact_governance_proposal:9e5a4eab5ea3c147310f6e6f8011c7cd` with material +`sha256:9e5a4eab5ea3c147310f6e6f8011c7cd56682933711e3be8924fbe294426b40d`. + +## Verification + +The frozen World dependency versions plus the stacked Core candidate and separately packaged +reference action adapter produced: + +```text +python -B -m pytest domain_packs/tests/test_p2c3_measured_feedback.py \ + domain_packs/tests/test_p2c4_reviewed_impact_disposition.py \ + domain_packs/tests/test_p2c5_citation_correctness_outcome.py \ + domain_packs/tests/test_p2c6_contradiction_attention_outcome.py \ + domain_packs/tests/test_p2c7_correction_detection_delay_outcome.py \ + domain_packs/tests/test_p2c8_correction_revision_stability_outcome.py -q --tb=short +20 passed in 6.50s + +python -B -m pytest -q --tb=short +103 passed in 20.04s + +python -B -m pytest adapters/federal_register_source/tests -q --tb=short +26 passed in 0.26s + +python -B -m pytest tests/test_release_contract.py -q --tb=short +7 passed in 0.05s + +# Installed public ace-core==0.5.0; no Core source checkout or reference action adapter. +python -B -m pytest -q --tb=short -rs +82 passed, 21 skipped in 14.24s + +ruff check --no-cache +PASS + +ruff format --check --no-cache +3 files already formatted + +uv build --out-dir +Successfully built unchanged 0.9.0 source distribution and inert data-only wheel + +git diff --check +PASS +``` + +The twenty-one public-Core skips are explicit boundaries: one P2C2 test requires the separately +packaged Core reference action adapter; P2C3-P2C8 require unreleased stacked Core candidate +contracts. The wheel contains 45 inert Domain Pack JSON/metadata files, no Python or entry points, +and retains `ace-core>=0.5.0,<0.6`. + +Repository-wide Ruff remains an inherited release-hygiene blocker. The exact locked check reports +the same 15 lint findings and 20 format targets as the P2C7 parent. Scoped P2C8 checks and +`git diff --check` are green. P2C8 does not rewrite unrelated history, but release closeout must +reconcile the repository-wide gate before publication. + +## Claim boundary + +The World reviewer exact-loads the prior and revised Briefs, then derives affected-update +correctness, correction visibility, complete source coverage, claim-count preservation, and the +exact stable-claim partition. The Core Outcome points to that review record. Historical replay +requires no new authority, and fresh hosts reproduce the exact Brief identities, claim sets, +classification, metrics, and proposal disposition. + +The paraphrase control is a frozen World policy fixture, not a general semantic-equivalence engine. +One correction and two replicated workflows do not establish live revision, a population stability +rate, calibration, source independence, general Brief quality, causality, legal truth, or human +benefit. The proposal remains non-effective and unapplied. + +## Remaining work + +Calibration, another independently sourced correction event, a materially different Market +journey, combined-main review/CI, public artifacts, repository-wide lint/format reconciliation, +security/release checks, and opt-in live transport remain future bounded work. Core issue #49 F1, +F3, and F5 still require explicit 0.6 release-owner disposition; this World packet neither +implements nor re-dates them. diff --git a/docs/design/world-intelligence-p2c8-correction-revision-stability-outcome-work-packet-v1.md b/docs/design/world-intelligence-p2c8-correction-revision-stability-outcome-work-packet-v1.md new file mode 100644 index 0000000..28d2642 --- /dev/null +++ b/docs/design/world-intelligence-p2c8-correction-revision-stability-outcome-work-packet-v1.md @@ -0,0 +1,110 @@ +# World Intelligence P2C8 correction-revision-stability outcome work packet (v1) + +**Status:** stacked source-checkout candidate; this packet does not release World Intelligence, +apply a governance proposal, close ACE Core issue #38, pass SI4, or complete ACE 0.6.0. + +**Frozen:** 2026-08-10 from World P2C7 commit +`216a37a1fdcb4f0baf9ac148ab9e525559141c22`, stacked on the Core exact observed-result +provenance candidate `433e3d16c5458c975557dcd1552824fb959d4d12`. + +## Objective + +Measure whether a correction-induced Brief revision changes the claim it should while preserving +the exact identities of claims the correction does not affect. The packet extends the governed +recorded-data journey: + +```text +original Observation + correction Observation -> prior Brief + -> treatment revision / unrelated-drift control + -> Decision -> reviewed Action -> exact independent revision review + -> observed Outcome -> useful / harmful / unproven evaluation -> proposal only +``` + +The source pair remains FCC Federal Register document `2020-28779` and its explicit correction +`2021-10670`. P2C8 creates actual domain-neutral `BriefV1Alpha1` records rather than a weaker +World-only summary shape. The prior Brief has one affected instruction claim and two stable facts. +Both revised Briefs replace the affected claim with the exact correction instruction, retain the +same two source citations, and keep the claim count at three. Treatment reuses the exact two stable +claim identities; the control paraphrases both stable facts and therefore changes their content +identities despite unchanged correction semantics and source coverage. + +## Product-owned review policy + +World owns `world_recorded_correction_revision_stability` version `candidate-1`. Its source fixture, +prior Brief, affected claim, replacement claim, stable claim set, reviewer, formula, and policy +digest are inspectable. The score is the unaffected-claim preservation rate, but only when all four +gates pass: + +1. the exact replacement claim is present and the stale affected claim is absent; +2. the correction citation and exact correction/prior lineage are visible; +3. original and correction source coverage is complete; and +4. the prior and revised claim counts match. + +If any gate fails, the score is `0`. Treatment scores `1.0`; the unrelated-drift control scores +`0.0`. Core and Intelligence see only exact Brief/result coordinates and a scalar outcome. FCC, +Federal Register, correction, affected/unaffected claim policy, paraphrase classification, and +review vocabulary remain in World. + +## Exact acceptance + +P2C8 must: + +1. rerun P2C2 through P2C7 and preserve every prior immutable record, evaluation, proposal, and + reviewed disposition; +2. append an exact prior `BriefV1Alpha1` citing the original Observation; +3. append treatment and control `BriefV1Alpha1` revisions with exact prior, original, and + correction lineage; +4. prove both revisions remove the stale affected claim, add the same exact replacement claim, + cite the same original/correction sources, and retain the same claim count; +5. prove treatment preserves both stable claim identities while control preserves neither and + introduces exactly two unrelated content identities; +6. create two distinct reviewed treatment/control Action pairs under matched task conditions; +7. append four independently authenticated review records and four Core Outcomes naming those + exact review records; +8. classify the bounded two-pair difference `useful`, emit only a non-effective `promote` + proposal, and perform no proposal application; +9. replay without reauthorization and reproduce Brief identities, claim sets, classification, and + substantive metrics across fresh hosts; and +10. reject duplicate claims, incomplete stable-claim partitions, missing correction visibility + paired with a positive score, and caller-invented aggregate scores. + +## Negative and failure controls + +The unrelated-drift Brief is the primary product negative control. It has the same exact prior, +source Observations, citations, claim count, replacement claim, reviewed workflow, and matched +conditions as treatment. Only the two unaffected claims are paraphrased, changing their content +identities. This is a frozen fixture classification, not a general semantic-equivalence engine. + +`BriefV1Alpha1` rejects duplicate claim identities. The review contract requires the preserved and +drifted sets to partition the expected stable claims exactly and derives its rate and score from +that partition plus the four gates. Exact review loading rejects unavailable, changed, relabelled, +cross-product, or incomplete Brief material. Stacked Core tests remain authoritative for missing +attribution/result provenance, condition mismatch, cutoff leakage, unavailable Outcomes, +duplicate/replayed evidence, interruption, restart, and denied authority. + +## Files and rollback + +This packet owns: + +- `scripts/p2c8_correction_revision_stability_outcome.py`; +- `domain_packs/tests/test_p2c8_correction_revision_stability_outcome.py`; +- the additive P2C7 state handoff; +- this work packet, its audit, and restrained README/roadmap references. + +It changes no shipped Domain Pack, connector, fixture source policy, package version, dependency +range, lockfile, release record, Core contract, or public artifact. Rollback removes the P2C8 +harness, tests, state handoff, and candidate documentation. Brief, review, Outcome, evaluation, +and proposal records already persisted by a host remain immutable history. + +## Non-claims and next packet + +This is one recorded correction pair and two replicated matched workflows. It establishes exact +claim-identity preservation and criterion sensitivity under one frozen World rule. It does not +establish live revision, general semantic equivalence, a population stability rate, calibration, +source independence, general Brief quality, causal benefit, legal truth, or human usefulness. + +The next bounded outcome packet should freeze calibration under a declared forecast/observed-result +rule or repeat the unchanged correction/revision contracts over a materially different source. +Independent Market reproduction, public Core artifacts, combined-main review/CI, compatibility, +security and release gates, opt-in live transport, issue #49 disposition, repository-wide hygiene, +and any separately authorized proposal application remain separate work. diff --git a/domain_packs/tests/test_p2c8_correction_revision_stability_outcome.py b/domain_packs/tests/test_p2c8_correction_revision_stability_outcome.py new file mode 100644 index 0000000..1abfd10 --- /dev/null +++ b/domain_packs/tests/test_p2c8_correction_revision_stability_outcome.py @@ -0,0 +1,157 @@ +from __future__ import annotations + +import copy +import importlib.util +import json + +import pytest +from pydantic import ValidationError + + +def _require_candidate_contracts() -> None: + if importlib.util.find_spec("ace.application.measured_impact") is None: + pytest.skip("P2C8 requires the stacked ACE Core measured-impact candidate") + from ace.intelligence import ImpactOutcomeMeasuresV1Alpha1 + + if "observed_result" not in ImpactOutcomeMeasuresV1Alpha1.model_fields: + pytest.skip("P2C8 requires exact observed-result provenance from the stacked Core candidate") + if importlib.util.find_spec("ace_reference_workspace_action") is None: + pytest.skip("P2C8 requires the separately packaged Core reference adapter") + + +@pytest.mark.asyncio +async def test_exact_brief_revision_review_becomes_a_measured_outcome(tmp_path) -> None: + _require_candidate_contracts() + from scripts.p2c8_correction_revision_stability_outcome import run_correction_revision_stability_outcome + + result = await run_correction_revision_stability_outcome(tmp_path) + + assert result["review_policy"]["reviewer_ref"] == "principal:world-revision-stability-reviewer" + assert result["evaluation"]["criterion"]["requires_observed_result"] is True + assert result["evaluation"]["classification"] == "useful" + assert result["evaluation"]["matched_pair_count"] == 2 + assert result["evaluation"]["mean_effect"] == 1.0 + assert result["proposal"]["action"] == "promote" + assert result["proposal"]["live_effect"] is False + assert result["proposal"]["selectable"] is False + assert result["proposal"]["requires_human_review"] is True + assert result["replay"]["historical"] is True + assert result["replay"]["no_reauthorization"] is True + assert {item["contract"] for item in result["briefs"].values()} == {"ace.intelligence.brief/v1alpha1"} + + +@pytest.mark.asyncio +async def test_equal_coverage_control_isolates_unaffected_claim_identity_stability(tmp_path) -> None: + _require_candidate_contracts() + from scripts.p2c8_correction_revision_stability_outcome import run_correction_revision_stability_outcome + + result = await run_correction_revision_stability_outcome(tmp_path) + prior = result["briefs"]["prior"] + treatment = result["briefs"]["treatment"] + control = result["briefs"]["control"] + expected = result["expected_revision"] + treatment_reviews = result["observed_results"]["treatment"] + control_reviews = result["observed_results"]["control"] + + assert len(prior["claims"]) == len(treatment["claims"]) == len(control["claims"]) == 3 + assert len(treatment["citations"]) == len(control["citations"]) == 2 + assert {item["affected_update_correct"] for item in (*treatment_reviews, *control_reviews)} == {True} + assert {item["correction_visible"] for item in (*treatment_reviews, *control_reviews)} == {True} + assert {item["source_coverage_complete"] for item in (*treatment_reviews, *control_reviews)} == {True} + assert {item["claim_count_preserved"] for item in (*treatment_reviews, *control_reviews)} == {True} + assert {tuple(item["preserved_stable_claim_ids"]) for item in treatment_reviews} == { + tuple(expected["stable_claim_ids"]) + } + assert {tuple(item["drifted_stable_claim_ids"]) for item in treatment_reviews} == {()} + assert {tuple(item["preserved_stable_claim_ids"]) for item in control_reviews} == {()} + assert {tuple(item["drifted_stable_claim_ids"]) for item in control_reviews} == { + tuple(expected["stable_claim_ids"]) + } + assert {item["revision_stability_score"] for item in treatment_reviews} == {1.0} + assert {item["revision_stability_score"] for item in control_reviews} == {0.0} + + +@pytest.mark.asyncio +async def test_revised_briefs_name_the_prior_and_exact_correction_without_rewriting_history(tmp_path) -> None: + _require_candidate_contracts() + from scripts.p2c8_correction_revision_stability_outcome import run_correction_revision_stability_outcome + + result = await run_correction_revision_stability_outcome(tmp_path) + prior = result["briefs"]["prior"] + expected = result["expected_revision"] + source_pair = result["source_pair"] + + for variant in ("treatment", "control"): + revised = result["briefs"][variant] + lineage_ids = {item["resource_id"] for item in revised["lineage"]} + claim_ids = {item["claim_id"] for item in revised["claims"]} + assert prior["resource_id"] in lineage_ids + assert source_pair["original_observation_id"] in lineage_ids + assert source_pair["correction_observation_id"] in lineage_ids + assert expected["affected_claim_id"] not in claim_ids + assert expected["replacement_claim_id"] in claim_ids + assert result["scope"]["actual_brief_contracts"] is True + assert result["scope"]["unaffected_claim_identity_preservation_reviewed"] is True + assert result["scope"]["live_revision_claimed"] is False + assert result["scope"]["semantic_equivalence_engine_claimed"] is False + assert result["scope"]["proposal_applied"] is False + + +@pytest.mark.asyncio +async def test_duplicate_claims_missing_correction_visibility_and_invented_scores_fail_closed(tmp_path) -> None: + _require_candidate_contracts() + from ace.intelligence import BriefV1Alpha1 + + from scripts.p2c8_correction_revision_stability_outcome import ( + BriefRevisionStabilityReviewV1Alpha1, + run_correction_revision_stability_outcome, + ) + + result = await run_correction_revision_stability_outcome(tmp_path) + duplicate = copy.deepcopy(result["briefs"]["treatment"]) + duplicate["claims"] = [duplicate["claims"][0], duplicate["claims"][0], duplicate["claims"][2]] + duplicate["resource_id"] = None + duplicate["resource_digest"] = None + with pytest.raises(ValidationError, match="unique content identities"): + BriefV1Alpha1.model_validate_json(json.dumps(duplicate)) + + review = copy.deepcopy(result["observed_results"]["treatment"][0]) + review["correction_visible"] = False + review["review_id"] = None + review["review_digest"] = None + with pytest.raises(ValidationError, match="score differs from frozen product rule"): + BriefRevisionStabilityReviewV1Alpha1.model_validate_json(json.dumps(review)) + + review = copy.deepcopy(result["observed_results"]["treatment"][0]) + review["preserved_stable_claim_ids"] = review["preserved_stable_claim_ids"][:1] + review["review_id"] = None + review["review_digest"] = None + with pytest.raises(ValidationError, match="exactly partition stable claims"): + BriefRevisionStabilityReviewV1Alpha1.model_validate_json(json.dumps(review)) + + +@pytest.mark.asyncio +async def test_packet_briefs_and_classification_are_deterministic_across_fresh_hosts(tmp_path) -> None: + _require_candidate_contracts() + from scripts.p2c8_correction_revision_stability_outcome import run_correction_revision_stability_outcome + + first_root = tmp_path / "first" + second_root = tmp_path / "second" + first_root.mkdir() + second_root.mkdir() + first = await run_correction_revision_stability_outcome(first_root) + second = await run_correction_revision_stability_outcome(second_root) + + assert first["source_pair"] == second["source_pair"] + assert first["expected_revision"] == second["expected_revision"] + for variant in ("prior", "treatment", "control"): + assert first["briefs"][variant]["resource_id"] == second["briefs"][variant]["resource_id"] + assert first["briefs"][variant]["resource_digest"] == second["briefs"][variant]["resource_digest"] + for variant in ("treatment", "control"): + first_scores = [item["revision_stability_score"] for item in first["observed_results"][variant]] + second_scores = [item["revision_stability_score"] for item in second["observed_results"][variant]] + assert first_scores == second_scores + for field in ("classification", "matched_pair_count", "mean_effect"): + assert first["evaluation"][field] == second["evaluation"][field] + for field in ("action", "live_effect", "selectable", "requires_human_review"): + assert first["proposal"][field] == second["proposal"][field] diff --git a/scripts/p2c7_correction_detection_delay_outcome.py b/scripts/p2c7_correction_detection_delay_outcome.py index 3cf2fb5..3c99b27 100644 --- a/scripts/p2c7_correction_detection_delay_outcome.py +++ b/scripts/p2c7_correction_detection_delay_outcome.py @@ -544,10 +544,14 @@ async def _record_review_outcome( ) -async def run_correction_detection_delay_outcome(workspace_root: Path) -> dict[str, Any]: +async def run_correction_detection_delay_outcome( + workspace_root: Path, + *, + state_sink: dict[str, Any] | None = None, +) -> dict[str, Any]: """Run P2C7 over an exact recorded official correction pair.""" - state: dict[str, Any] = {} + state: dict[str, Any] = {} if state_sink is None else state_sink prior = await run_contradiction_attention_outcome(workspace_root, state_sink=state) environment = state["environment"] fixture = load_correction_fixture() @@ -746,6 +750,20 @@ async def run_correction_detection_delay_outcome(workspace_root: Path) -> dict[s if {item.detection_delay_seconds for item in control_reviews} != {21_600}: raise AssertionError("control did not preserve the exact six-hour replay delay") + state.update( + { + "p2c7_fixture": fixture, + "p2c7_original_observation": original, + "p2c7_original_observation_ref": original_ref, + "p2c7_correction_observation": correction, + "p2c7_correction_observation_ref": correction_ref, + "p2c7_treatment_artifact": treatment_artifact, + "p2c7_treatment_artifact_ref": treatment_ref, + "p2c7_control_artifact": control_artifact, + "p2c7_control_artifact_ref": control_ref, + } + ) + return { "contract": "ace.world-intelligence.correction-detection-delay-outcome/v1alpha1", "prior_contradiction_attention": prior, diff --git a/scripts/p2c8_correction_revision_stability_outcome.py b/scripts/p2c8_correction_revision_stability_outcome.py new file mode 100644 index 0000000..90a2e32 --- /dev/null +++ b/scripts/p2c8_correction_revision_stability_outcome.py @@ -0,0 +1,844 @@ +"""Measure bounded correction-induced Brief revision stability.""" + +from __future__ import annotations + +import asyncio +import json +from datetime import datetime +from pathlib import Path +from typing import Any, Literal, Self + +from ace.application import MeasuredImpactService +from ace.core import ( + AuthenticatedRuntimeContextV1Alpha1, + CapabilityArtifactIdentityV1Alpha1, + GovernedOperationBindingV1Alpha1, + GovernedStateHeadPreconditionV1Alpha1, + ImmutableRecordReferenceV1, + OutcomeIntentV1Alpha1, + OutcomeV1Alpha1, + canonical_json, + capability_state_ref_for_artifact, +) +from ace.intelligence import ( + BriefV1Alpha1, + CitationV1Alpha1, + ClaimGroundingKind, + GroundedClaimV1Alpha1, + ImpactClassification, + ImpactConditionsV1Alpha1, + ImpactCriterionV1Alpha1, + ImpactEvaluationRequestV1Alpha1, + ImpactEvidenceV1Alpha1, + ImpactGovernanceAction, + ImpactMetricDirection, + ImpactOutcomeMeasuresV1Alpha1, + ImpactTargetKind, + IntelligenceResourceMode, + LineageReferenceV1Alpha1, + LineageRelation, + LineageResourceKind, + ObservationV1Alpha1, +) +from pydantic import Field, model_validator + +from scripts.p2c2_federal_register_monitor import _time +from scripts.p2c2_governed_reality_brief import _context, _head +from scripts.p2c3_measured_feedback import ( + ReviewedExport, + _append_value, + _authorize_append, + _digest, + _record_attribution, + _run_reviewed_export, +) +from scripts.p2c5_citation_correctness_outcome import _derive_identity, _FrozenModel +from scripts.p2c7_correction_detection_delay_outcome import ( + correction_fixture_digest, + run_correction_detection_delay_outcome, +) + +OUTCOME_TYPE = "independent_artifact_review" +MEASURE_ID = "recorded_correction_revision_stability" +CRITERION_ID = "impact_criterion:world-recorded-correction-revision-stability" +PRIOR_BRIEF_AT = _time("2026-08-10T23:14:00Z") +REVISED_BRIEF_AT = _time("2026-08-10T23:15:00Z") +CRITERION_FROZEN_AT = _time("2026-08-10T23:16:00Z") +REVIEW_POLICY_ID = "world_recorded_correction_revision_stability" +REVIEW_POLICY_VERSION = "candidate-1" + +IMPACT_ARTIFACT = CapabilityArtifactIdentityV1Alpha1( + capability="measured_impact_evaluation", + contract="ace.application.measured-impact-service/v1alpha1", + implementation_id="world_recorded_correction_revision_stability_candidate", + implementation_version="0.1.0", + artifact_digest="sha256:" + "b" * 64, +) + + +class BriefRevisionStabilityReviewV1Alpha1(_FrozenModel): + """Exact product review of one prior/revised Brief pair.""" + + contract: Literal["ace.world-intelligence.brief-revision-stability-review/v1alpha1"] = ( + "ace.world-intelligence.brief-revision-stability-review/v1alpha1" + ) + product_id: str + review_key: str + prior_brief: ImmutableRecordReferenceV1 + revised_brief: ImmutableRecordReferenceV1 + original_observation: ImmutableRecordReferenceV1 + correction_observation: ImmutableRecordReferenceV1 + reviewer_context: AuthenticatedRuntimeContextV1Alpha1 + policy_id: str + policy_version: str + policy_digest: str + source_fixture_digest: str + expected_affected_claim_id: str + expected_replacement_claim_id: str + expected_stable_claim_ids: tuple[str, ...] = Field(min_length=1) + preserved_stable_claim_ids: tuple[str, ...] + drifted_stable_claim_ids: tuple[str, ...] + unexpected_claim_ids: tuple[str, ...] + prior_claim_count: int = Field(gt=0) + revised_claim_count: int = Field(gt=0) + replacement_claim_present: bool + stale_affected_claim_present: bool + affected_update_correct: bool + correction_visible: bool + source_coverage_complete: bool + claim_count_preserved: bool + unaffected_preservation_rate: float = Field(ge=0.0, le=1.0) + revision_stability_score: float = Field(ge=0.0, le=1.0) + limitations: tuple[str, ...] + reviewed_at: datetime + review_id: str | None = None + review_digest: str | None = None + + @model_validator(mode="after") + def validate_partition_score_and_identity(self) -> Self: + tuples = ( + self.expected_stable_claim_ids, + self.preserved_stable_claim_ids, + self.drifted_stable_claim_ids, + self.unexpected_claim_ids, + ) + if any(items != tuple(sorted(set(items))) for items in tuples): + raise ValueError("revision review claim identities must be unique and sorted") + expected = set(self.expected_stable_claim_ids) + preserved = set(self.preserved_stable_claim_ids) + drifted = set(self.drifted_stable_claim_ids) + if preserved & drifted or preserved | drifted != expected: + raise ValueError("preserved and drifted claims must exactly partition stable claims") + if self.expected_affected_claim_id in expected: + raise ValueError("the affected claim cannot also be an expected stable claim") + expected_update = self.replacement_claim_present and not self.stale_affected_claim_present + if self.affected_update_correct != expected_update: + raise ValueError("affected update disposition differs from exact claim presence") + expected_count_preserved = self.prior_claim_count == self.revised_claim_count + if self.claim_count_preserved != expected_count_preserved: + raise ValueError("claim-count disposition differs from exact counts") + expected_rate = len(preserved) / len(expected) + if self.unaffected_preservation_rate != expected_rate: + raise ValueError("unaffected preservation rate differs from exact claim partition") + gates = ( + self.affected_update_correct, + self.correction_visible, + self.source_coverage_complete, + self.claim_count_preserved, + ) + expected_score = expected_rate if all(gates) else 0.0 + if self.revision_stability_score != expected_score: + raise ValueError("revision stability score differs from frozen product rule") + _derive_identity( + self, + prefix="brief_revision_stability_review", + id_field="review_id", + digest_field="review_digest", + ) + return self + + +class _ReplayMustNotAuthorize: + async def authorize_action(self, request): + raise AssertionError( + f"historical revision-stability evaluation requested new authority: {request.authorization_key}" + ) + + +def _lineage( + resource: ObservationV1Alpha1 | BriefV1Alpha1, + *, + relation: LineageRelation = LineageRelation.DERIVED_FROM, +) -> LineageReferenceV1Alpha1: + if isinstance(resource, ObservationV1Alpha1): + kind = LineageResourceKind.OBSERVATION + available_at = resource.ingested_at + else: + kind = LineageResourceKind.BRIEF + available_at = resource.generated_at + return LineageReferenceV1Alpha1( + resource_kind=kind, + relation=relation, + resource_id=str(resource.resource_id), + resource_digest=str(resource.resource_digest), + resource_as_of=resource.as_of, + resource_available_at=available_at, + ) + + +def _citation( + observation: ObservationV1Alpha1, + *, + locator: str, + excerpt: str, +) -> CitationV1Alpha1: + return CitationV1Alpha1( + source_ref=observation.source_ref, + source_digest=observation.source_digest, + acquisition_mode=observation.acquisition_mode, + acquisition_receipt_ref=observation.acquisition_receipt_ref, + acquisition_receipt_digest=observation.acquisition_receipt_digest, + source_as_of=observation.source_published_at or observation.observed_at, + retrieved_at=observation.ingested_at, + locator=locator, + excerpt=excerpt, + ) + + +def _claim(statement: str, citation: CitationV1Alpha1) -> GroundedClaimV1Alpha1: + return GroundedClaimV1Alpha1( + statement=statement, + grounding_kind=ClaimGroundingKind.CITED, + citation_ids=(str(citation.citation_id),), + confidence=1.0, + uncertainty="Bounded to the exact recorded Federal Register documents and source-policy limits.", + ) + + +def _body(title: str, claims: tuple[GroundedClaimV1Alpha1, ...]) -> str: + return "\n".join((f"# {title}", "", *(f"- {item.statement}" for item in claims))) + "\n" + + +def _build_briefs(state: dict[str, Any]) -> dict[str, Any]: + environment = state["environment"] + original: ObservationV1Alpha1 = state["p2c7_original_observation"] + correction: ObservationV1Alpha1 = state["p2c7_correction_observation"] + activation_revision = state["brief_admission"].brief.activation_revision + original_citation = _citation( + original, + locator="Federal Register 85 FR 85524 and page 85530 instruction material", + excerpt="FCC document 2020-28779, 85 FR 85524.", + ) + correction_citation = _citation( + correction, + locator="Federal Register 86 FR 27275 correction to page 85530", + excerpt=("Remove instruction 20a and redesignate instructions 20b and 20c as instructions 20a and 20b."), + ) + affected = _claim( + "The page 85530 amendment instructions include instructions 20a, 20b, and 20c.", + original_citation, + ) + stable_publication = _claim( + "FCC document 2020-28779 was published on 2020-12-29 at 85 FR 85524.", + original_citation, + ) + stable_agency = _claim( + "The issuing agency is the Federal Communications Commission.", + original_citation, + ) + prior_claims = (affected, stable_publication, stable_agency) + prior = BriefV1Alpha1( + product_id=environment.fixture["product_id"], + mode=IntelligenceResourceMode.PREPARED, + activation_revision=activation_revision, + as_of=PRIOR_BRIEF_AT, + lineage=(_lineage(original),), + brief_type_ref="brief_type:world-reality-brief", + title="FCC Electronic Filing Rule — Recorded Brief", + executive_summary="Recorded orientation to the original FCC electronic-filing rule instructions.", + body_markdown=_body("FCC Electronic Filing Rule — Recorded Brief", prior_claims), + generated_at=PRIOR_BRIEF_AT, + citations=(original_citation,), + claims=prior_claims, + ) + replacement = _claim( + ( + "Correction 2021-10670 directs removal of instruction 20a and redesignation of " + "instructions 20b and 20c as instructions 20a and 20b on page 85530." + ), + correction_citation, + ) + treatment_claims = (replacement, stable_publication, stable_agency) + control_publication = _claim( + "Publication of FCC document 2020-28779 occurred on 2020-12-29 in 85 FR 85524.", + original_citation, + ) + control_agency = _claim( + "The Federal Communications Commission issued the document.", + original_citation, + ) + control_claims = (replacement, control_publication, control_agency) + common = { + "product_id": environment.fixture["product_id"], + "mode": IntelligenceResourceMode.PREPARED, + "activation_revision": activation_revision, + "as_of": REVISED_BRIEF_AT, + "lineage": ( + _lineage(original), + _lineage(correction), + _lineage(prior, relation=LineageRelation.CONTEXT), + ), + "brief_type_ref": "brief_type:world-reality-brief", + "title": "FCC Electronic Filing Rule — Corrected Brief", + "executive_summary": "The explicit FCC correction is visible while unaffected facts remain stable.", + "generated_at": REVISED_BRIEF_AT, + "citations": (original_citation, correction_citation), + } + treatment = BriefV1Alpha1( + **common, + body_markdown=_body("FCC Electronic Filing Rule — Corrected Brief", treatment_claims), + claims=treatment_claims, + ) + control = BriefV1Alpha1( + **common, + body_markdown=_body("FCC Electronic Filing Rule — Corrected Brief", control_claims), + claims=control_claims, + ) + return { + "prior": prior, + "treatment": treatment, + "control": control, + "original_citation": original_citation, + "correction_citation": correction_citation, + "affected_claim": affected, + "replacement_claim": replacement, + "stable_claims": (stable_publication, stable_agency), + } + + +async def _append_brief( + state: dict[str, Any], + *, + brief: BriefV1Alpha1, + role: str, +) -> ImmutableRecordReferenceV1: + environment = state["environment"] + requested_at = state["clock"]() + authorization = await _authorize_append( + state, + context=environment.context, + authorization_key=f"append:world-correction-revision-brief:{role}", + subject_ref=str(brief.resource_id), + subject_digest=str(brief.resource_digest), + requested_at=requested_at, + ) + return await _append_value( + state, + value=brief, + record_kind="brief", + record_key=str(brief.resource_id), + transaction_key=f"world-correction-revision-brief:{role}:{brief.resource_id}", + as_of=brief.as_of, + authorization=authorization, + ) + + +async def _load_brief(state: dict[str, Any], reference: ImmutableRecordReferenceV1) -> BriefV1Alpha1: + record = await state["environment"].store.load_record( + reference.storage_id, + product_id=reference.product_id, + record_space=reference.record_space, + record_kind=reference.record_kind, + ) + if ( + record is None + or record.reference() != reference + or record.payload_contract != "ace.intelligence.brief/v1alpha1" + ): + raise AssertionError("revision-stability Brief is unavailable or changed") + return BriefV1Alpha1.model_validate(record.payload) + + +def _policy_digest( + *, + fixture_digest: str, + prior_ref: ImmutableRecordReferenceV1, + affected_claim_id: str, + replacement_claim_id: str, + stable_claim_ids: tuple[str, ...], +) -> str: + return _digest( + { + "policy_id": REVIEW_POLICY_ID, + "policy_version": REVIEW_POLICY_VERSION, + "fixture_digest": fixture_digest, + "prior_brief": prior_ref.model_dump(mode="json"), + "expected_affected_claim_id": affected_claim_id, + "expected_replacement_claim_id": replacement_claim_id, + "expected_stable_claim_ids": stable_claim_ids, + "score": ( + "unaffected preservation rate when the affected update, correction visibility, " + "source coverage, and claim count all pass; otherwise 0" + ), + } + ) + + +async def _review_revision( + state: dict[str, Any], + *, + material: dict[str, Any], + prior_ref: ImmutableRecordReferenceV1, + revised_ref: ImmutableRecordReferenceV1, + pair_index: int, + variant: str, +) -> tuple[BriefRevisionStabilityReviewV1Alpha1, ImmutableRecordReferenceV1]: + environment = state["environment"] + prior = await _load_brief(state, prior_ref) + revised = await _load_brief(state, revised_ref) + expected_stable = tuple(sorted(str(item.claim_id) for item in material["stable_claims"])) + affected_id = str(material["affected_claim"].claim_id) + replacement_id = str(material["replacement_claim"].claim_id) + prior_ids = {str(item.claim_id) for item in prior.claims} + revised_ids = {str(item.claim_id) for item in revised.claims} + preserved = tuple(sorted(set(expected_stable) & revised_ids)) + drifted = tuple(sorted(set(expected_stable) - revised_ids)) + unexpected = tuple(sorted(revised_ids - set(expected_stable) - {replacement_id})) + correction_citation_id = str(material["correction_citation"].citation_id) + original_citation_id = str(material["original_citation"].citation_id) + revised_citation_ids = {str(item.citation_id) for item in revised.citations} + lineage = {(item.resource_id, item.resource_digest) for item in revised.lineage} + correction = state["p2c7_correction_observation"] + prior_lineage = (str(prior.resource_id), str(prior.resource_digest)) in lineage + correction_lineage = (str(correction.resource_id), str(correction.resource_digest)) in lineage + replacement_present = replacement_id in revised_ids + stale_present = affected_id in revised_ids + correction_visible = correction_citation_id in revised_citation_ids and correction_lineage and prior_lineage + source_coverage_complete = {original_citation_id, correction_citation_id}.issubset(revised_citation_ids) + claim_count_preserved = len(prior.claims) == len(revised.claims) + preservation_rate = len(preserved) / len(expected_stable) + affected_update_correct = replacement_present and not stale_present + score = ( + preservation_rate + if affected_update_correct and correction_visible and source_coverage_complete and claim_count_preserved + else 0.0 + ) + reviewed_at = state["clock"]() + reviewer = _context(environment.context, "principal:world-revision-stability-reviewer") + fixture_digest = correction_fixture_digest(state["p2c7_fixture"]) + review = BriefRevisionStabilityReviewV1Alpha1( + product_id=environment.fixture["product_id"], + review_key=f"recorded-correction-revision-stability:{pair_index}:{variant}", + prior_brief=prior_ref, + revised_brief=revised_ref, + original_observation=state["p2c7_original_observation_ref"], + correction_observation=state["p2c7_correction_observation_ref"], + reviewer_context=reviewer, + policy_id=REVIEW_POLICY_ID, + policy_version=REVIEW_POLICY_VERSION, + policy_digest=_policy_digest( + fixture_digest=fixture_digest, + prior_ref=prior_ref, + affected_claim_id=affected_id, + replacement_claim_id=replacement_id, + stable_claim_ids=expected_stable, + ), + source_fixture_digest=fixture_digest, + expected_affected_claim_id=affected_id, + expected_replacement_claim_id=replacement_id, + expected_stable_claim_ids=expected_stable, + preserved_stable_claim_ids=preserved, + drifted_stable_claim_ids=drifted, + unexpected_claim_ids=unexpected, + prior_claim_count=len(prior.claims), + revised_claim_count=len(revised.claims), + replacement_claim_present=replacement_present, + stale_affected_claim_present=stale_present, + affected_update_correct=affected_update_correct, + correction_visible=correction_visible, + source_coverage_complete=source_coverage_complete, + claim_count_preserved=claim_count_preserved, + unaffected_preservation_rate=preservation_rate, + revision_stability_score=score, + limitations=( + "recorded_replay_not_live_revision", + "one_explicit_correction_pair", + "two_replicated_workflows_not_independent_events", + "semantic_equivalence_of_paraphrases_is_product_fixture_policy", + ), + reviewed_at=reviewed_at, + ) + authorization = await _authorize_append( + state, + context=reviewer, + authorization_key=f"recorded-correction-revision-stability:{pair_index}:{variant}", + subject_ref=str(review.review_id), + subject_digest=str(review.review_digest), + requested_at=reviewed_at, + ) + reference = await _append_value( + state, + value=review, + record_kind="brief_revision_stability_review", + record_key=str(review.review_id), + transaction_key=f"brief-revision-stability-review:{review.review_id}", + as_of=reviewed_at, + authorization=authorization, + ) + if affected_id not in prior_ids: + raise AssertionError("frozen affected claim is absent from the exact prior Brief") + return review, reference + + +async def _record_review_outcome( + state: dict[str, Any], + *, + export: ReviewedExport, + review: BriefRevisionStabilityReviewV1Alpha1, + review_ref: ImmutableRecordReferenceV1, + pair_index: int, + variant: str, +) -> ImmutableRecordReferenceV1: + environment = state["environment"] + measures = ImpactOutcomeMeasuresV1Alpha1( + primary_value=review.revision_stability_score, + observed_result=review_ref, + cost_usd=0.0, + failure_count=0, + degraded=False, + limitations=review.limitations, + ) + recorded_at = state["clock"]() + observer = _context(environment.context, "principal:world-revision-stability-outcome-observer") + intent = OutcomeIntentV1Alpha1( + product_id=environment.fixture["product_id"], + authenticated_context=observer, + decision=export.decision_ref, + outcome_type=OUTCOME_TYPE, + measure_id=MEASURE_ID, + value_json=canonical_json(measures.model_dump(mode="json")), + observed_at=review.reviewed_at, + recorded_at=recorded_at, + ) + authorization = await _authorize_append( + state, + context=observer, + authorization_key=f"recorded-correction-revision-stability-outcome:{pair_index}:{variant}", + subject_ref=str(intent.intent_id), + subject_digest=str(intent.intent_digest), + requested_at=recorded_at, + ) + outcome = OutcomeV1Alpha1(intent=intent, authorization=authorization) + return await _append_value( + state, + value=outcome, + record_kind="outcome", + record_key=str(outcome.outcome_id), + transaction_key=f"outcome:{outcome.outcome_id}", + as_of=review.reviewed_at, + authorization=authorization, + ) + + +def _install_policy(state: dict[str, Any]): + environment = state["environment"] + runtime = state["runtime"] + product_id = environment.fixture["product_id"] + criterion_head = _head(product_id, "impact_criterion", CRITERION_ID, 90) + operation_head = _head( + product_id, + "governed_operation_configuration", + "governed_operation_configuration:world-recorded-correction-revision-stability", + 91, + ) + binding = GovernedOperationBindingV1Alpha1( + product_id=product_id, + artifact=IMPACT_ARTIFACT, + configuration_ref=operation_head.state_id, + authority="append_measured_impact", + grant_ref="authority_grant:world-recorded-correction-revision-stability", + state_head_precondition=GovernedStateHeadPreconditionV1Alpha1.from_head(operation_head), + ) + capability_head = _head( + product_id, + "capability_state", + capability_state_ref_for_artifact(IMPACT_ARTIFACT), + 92, + ) + authority_head = _head(product_id, "authority_grant", binding.grant_ref, 93) + for head in (criterion_head, operation_head, capability_head, authority_head): + environment.store.set_governed_state_head(head) + runtime.heads[head.state_kind, head.state_id] = head + runtime.bindings = (*runtime.bindings, binding) + return criterion_head, binding + + +async def run_correction_revision_stability_outcome(workspace_root: Path) -> dict[str, Any]: + """Run P2C8 over exact prior, stable revision, and drift-control Briefs.""" + + state: dict[str, Any] = {} + prior_packet = await run_correction_detection_delay_outcome(workspace_root, state_sink=state) + environment = state["environment"] + material = _build_briefs(state) + prior_ref = await _append_brief(state, brief=material["prior"], role="prior") + treatment_ref = await _append_brief(state, brief=material["treatment"], role="treatment") + control_ref = await _append_brief(state, brief=material["control"], role="drift-control") + treatment_exports = tuple( + [ + await _run_reviewed_export( + state, + subject=treatment_ref, + content=material["treatment"].body_markdown, + pair_index=index, + variant="correction-revision-treatment", + ) + for index in (1, 2) + ] + ) + control_exports = tuple( + [ + await _run_reviewed_export( + state, + subject=control_ref, + content=material["control"].body_markdown, + pair_index=index, + variant="correction-revision-control", + ) + for index in (1, 2) + ] + ) + criterion_head, impact_binding = _install_policy(state) + criterion = ImpactCriterionV1Alpha1( + product_id=environment.fixture["product_id"], + criterion_id=CRITERION_ID, + criterion_version="candidate-1", + target_kind=ImpactTargetKind.INTELLIGENCE_ARTIFACT, + outcome_type=OUTCOME_TYPE, + measure_id=MEASURE_ID, + metric_direction=ImpactMetricDirection.HIGHER_IS_BETTER, + useful_effect_threshold=0.75, + harmful_effect_threshold=0.75, + minimum_matched_pairs=2, + requires_observed_result=True, + harmful_action=ImpactGovernanceAction.ROLLBACK, + state_head_precondition=GovernedStateHeadPreconditionV1Alpha1.from_head(criterion_head), + frozen_at=CRITERION_FROZEN_AT, + ) + + evidence: list[ImpactEvidenceV1Alpha1] = [] + treatment_reviews: list[BriefRevisionStabilityReviewV1Alpha1] = [] + control_reviews: list[BriefRevisionStabilityReviewV1Alpha1] = [] + observed_times: list[datetime] = [] + for index, (treatment_export, control_export) in enumerate( + zip(treatment_exports, control_exports, strict=True), start=1 + ): + treatment_attribution = await _record_attribution( + state, + export=treatment_export, + subject=treatment_ref, + pair_index=index, + variant="correction-revision-treatment", + ) + control_attribution = await _record_attribution( + state, + export=control_export, + subject=control_ref, + pair_index=index, + variant="correction-revision-control", + ) + treatment_review, treatment_review_ref = await _review_revision( + state, + material=material, + prior_ref=prior_ref, + revised_ref=treatment_ref, + pair_index=index, + variant="treatment", + ) + control_review, control_review_ref = await _review_revision( + state, + material=material, + prior_ref=prior_ref, + revised_ref=control_ref, + pair_index=index, + variant="control", + ) + treatment_outcome = await _record_review_outcome( + state, + export=treatment_export, + review=treatment_review, + review_ref=treatment_review_ref, + pair_index=index, + variant="treatment", + ) + control_outcome = await _record_review_outcome( + state, + export=control_export, + review=control_review, + review_ref=control_review_ref, + pair_index=index, + variant="control", + ) + treatment_reviews.append(treatment_review) + control_reviews.append(control_review) + observed_times.extend((treatment_review.reviewed_at, control_review.reviewed_at)) + conditions = ImpactConditionsV1Alpha1( + product_id=environment.fixture["product_id"], + condition_key=f"world-recorded-correction-revision-stability-pair:{index}", + route_id="world:fcc-recorded-correction-revision-review", + context_json=canonical_json( + { + "fixture_digest": correction_fixture_digest(state["p2c7_fixture"]), + "pair_index": index, + "prior_brief": prior_ref.model_dump(mode="json"), + "review_policy_digest": treatment_review.policy_digest, + "task": "exact_correction_update_with_unaffected_claim_identity_preservation", + } + ), + observation_window_start=CRITERION_FROZEN_AT, + observation_window_end=max(observed_times), + frozen_at=CRITERION_FROZEN_AT, + ) + evidence.append( + ImpactEvidenceV1Alpha1( + product_id=environment.fixture["product_id"], + evidence_key=f"world-recorded-correction-revision-stability:{index}", + treatment_attribution=treatment_attribution, + control_attribution=control_attribution, + treatment_decision=treatment_export.decision_ref, + control_decision=control_export.decision_ref, + treatment_action_review=treatment_export.review_ref, + treatment_action_admission=treatment_export.admission_ref, + treatment_action_terminal=treatment_export.terminal_ref, + control_action_review=control_export.review_ref, + control_action_admission=control_export.admission_ref, + control_action_terminal=control_export.terminal_ref, + treatment_outcome=treatment_outcome, + control_outcome=control_outcome, + treatment_conditions=conditions, + control_conditions=conditions, + ) + ) + + request = ImpactEvaluationRequestV1Alpha1( + evaluation_key="world-recorded-correction-revision-stability:fcc-2021-10670", + product_id=environment.fixture["product_id"], + authenticated_context=environment.context, + criterion=criterion, + target=treatment_ref, + control=control_ref, + evidence=tuple(evidence), + cutoff_at=state["clock"](), + requested_at=state["clock"](), + ) + service = MeasuredImpactService( + store=environment.store, + authorizer=state["reasoning"], + operation_binding=impact_binding, + ) + admission = await service.evaluate(request) + historical = await MeasuredImpactService( + store=environment.store, + authorizer=_ReplayMustNotAuthorize(), + operation_binding=impact_binding, + ).evaluate(request) + if admission.replayed or not historical.replayed or historical.evaluation != admission.evaluation: + raise AssertionError("revision-stability evaluation did not replay exact historical material") + if admission.evaluation.classification is not ImpactClassification.USEFUL: + raise AssertionError("frozen revision-stability criterion did not classify useful") + if admission.proposal is None or admission.proposal.action is not ImpactGovernanceAction.PROMOTE: + raise AssertionError("revision-stability result did not emit its proposal-only mapping") + if {item.revision_stability_score for item in treatment_reviews} != {1.0}: + raise AssertionError("treatment did not preserve every unaffected claim identity") + if {item.revision_stability_score for item in control_reviews} != {0.0}: + raise AssertionError("drift control did not expose gratuitous unrelated revision") + if {item.affected_update_correct for item in (*treatment_reviews, *control_reviews)} != {True}: + raise AssertionError("one revision lost the exact correction update") + if {item.source_coverage_complete for item in (*treatment_reviews, *control_reviews)} != {True}: + raise AssertionError("one revision changed exact source coverage") + + state.update( + { + "p2c8_prior_brief": material["prior"], + "p2c8_prior_brief_ref": prior_ref, + "p2c8_treatment_brief": material["treatment"], + "p2c8_treatment_brief_ref": treatment_ref, + "p2c8_control_brief": material["control"], + "p2c8_control_brief_ref": control_ref, + } + ) + return { + "contract": "ace.world-intelligence.correction-revision-stability-outcome/v1alpha1", + "prior_correction_detection": prior_packet, + "source_pair": { + "fixture_id": state["p2c7_fixture"]["fixture_id"], + "fixture_digest": correction_fixture_digest(state["p2c7_fixture"]), + "original_observation_id": str(state["p2c7_original_observation"].resource_id), + "correction_observation_id": str(state["p2c7_correction_observation"].resource_id), + }, + "review_policy": { + "policy_id": REVIEW_POLICY_ID, + "policy_version": REVIEW_POLICY_VERSION, + "policy_digest": treatment_reviews[0].policy_digest, + "reviewer_ref": treatment_reviews[0].reviewer_context.actor_ref, + "score_rule": ( + "unaffected preservation rate gated by exact correction update, correction visibility, " + "source coverage, and claim-count preservation" + ), + }, + "briefs": { + "prior": material["prior"].model_dump(mode="json"), + "treatment": material["treatment"].model_dump(mode="json"), + "control": material["control"].model_dump(mode="json"), + }, + "expected_revision": { + "affected_claim_id": str(material["affected_claim"].claim_id), + "replacement_claim_id": str(material["replacement_claim"].claim_id), + "stable_claim_ids": tuple(sorted(str(item.claim_id) for item in material["stable_claims"])), + }, + "observed_results": { + "treatment": tuple(item.model_dump(mode="json") for item in treatment_reviews), + "control": tuple(item.model_dump(mode="json") for item in control_reviews), + }, + "evaluation": admission.evaluation.model_dump(mode="json"), + "proposal": admission.proposal.model_dump(mode="json"), + "replay": { + "historical": historical.replayed, + "no_reauthorization": True, + "transaction_receipt_id": str(historical.transaction_receipt.receipt_id), + }, + "scope": { + "actual_brief_contracts": True, + "exact_correction_update_reviewed": True, + "unaffected_claim_identity_preservation_reviewed": True, + "equal_source_coverage_control": True, + "equal_claim_count_control": True, + "network_access": False, + "live_revision_claimed": False, + "semantic_equivalence_engine_claimed": False, + "population_revision_stability_claimed": False, + "general_brief_quality_claimed": False, + "human_benefit_claimed": False, + "causality_claimed": False, + "proposal_applied": False, + "autonomous_publication": False, + }, + } + + +def main() -> None: + import argparse + + parser = argparse.ArgumentParser() + parser.add_argument("workspace_root", type=Path) + args = parser.parse_args() + print( + json.dumps( + asyncio.run(run_correction_revision_stability_outcome(args.workspace_root)), + indent=2, + sort_keys=True, + ) + ) + + +if __name__ == "__main__": + main()