diff --git a/README.md b/README.md index 4cd5c90..0114271 100644 --- a/README.md +++ b/README.md @@ -201,6 +201,20 @@ because structural coverage does not establish citation correctness, general Bri benefit, causality, or live freshness. Exact replay returns the historical Decision without new authorization. This is explicit disposition, not reclassification or proposal application. +P2C5 adds a distinct product-owned outcome rather than broadening the structural score. The named +principal `principal:world-citation-correctness-reviewer` records an exact immutable review over +the Brief, its one cited claim, the exact two citation identities, and the two admitted official +Observation references. Each Core Outcome names that exact result. A matched negative control +retains both citation identities but swaps the two publication dates in the claim: treatment and +control both have `1.0` coverage, while correctness is `1.0` and `0.0` respectively across two +pairs. The domain-neutral evaluator classifies that bounded difference as useful and still emits +only a non-effective proposal requiring separate review. + +This candidate establishes exact review provenance and sensitivity to one semantic corruption. It +does not establish reviewer infallibility, current network freshness, source independence, general +Brief quality, causal impact, or human benefit. Citation review vocabulary and policy remain in +World; Core and Intelligence receive only the exact immutable observed-result coordinate. + ## What the public World proof demonstrates Generate a self-contained visual Reality Brief and its exact machine-readable backing data: @@ -356,10 +370,11 @@ workspace export followed by separate verification and promotion. The [World Intelligence roadmap](ROADMAP.md) owns current domain direction. Detailed packet history remains in [`docs/world-intelligence-roadmap-status-2026-08-06.md`](docs/world-intelligence-roadmap-status-2026-08-06.md), -and release history is in [`CHANGELOG.md`](CHANGELOG.md). P2C4 now demonstrates a separately -authorized reject/no-action disposition of the P2C3 proposal without effective state change. The -next bounded measurement work is an independently reviewed product outcome such as citation -correctness, contradiction coverage, correction quality, detection delay, or false-alert rate. +and release history is in [`CHANGELOG.md`](CHANGELOG.md). P2C4 demonstrates a separately +authorized reject/no-action disposition of the P2C3 proposal without effective state change. P2C5 +adds an independently reviewed citation-correctness Outcome and a citation-preserving semantic +negative control. The next bounded measurement work is contradiction/correction coverage, +detection delay, false-alert rate, or independent Market reproduction. Separately reviewed opt-in network transport and P2D multi-source conflict/correction with LIVE inputs remain independent work. None of these steps may add autonomous publishing, delivery, persuasion, or action authority to a Domain Pack. diff --git a/ROADMAP.md b/ROADMAP.md index c674e21..b181808 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -16,7 +16,7 @@ source code into the platform. See the [0.9.0 release record](docs/releases/world-intelligence-p2c2-v0.9.0.md). -## Candidate — P2C3/P2C4 measured feedback and reviewed disposition +## Candidate — P2C3–P2C5 measured feedback, disposition, and citation correctness - The exact 0.9.0 Brief and a World-owned source-only control pass through two matched reviewed export pairs under one frozen structural citation-coverage criterion. @@ -27,12 +27,18 @@ See the [0.9.0 release record](docs/releases/world-intelligence-p2c2-v0.9.0.md). - A separate authenticated and authorized World reviewer rejects broader promotion with an exact `no_action` Core Decision. The useful evaluation and promote proposal remain unchanged, and no effective governed-state head changes. -- This is source-checkout evidence against Core PR #88, not a released World capability, human +- A second frozen criterion requires exact independent review provenance behind every Outcome + score. A citation-preserving date-swap control retains `1.0` citation coverage while correctness + falls from `1.0` to `0.0` over two matched pairs; the bounded evaluation again emits only a + non-effective proposal. +- This is source-checkout evidence against stacked Core candidates, not a released World capability, human benefit finding, causal claim, network-freshness proof, or applied governance change. See the [P2C3 candidate work packet](docs/design/world-intelligence-p2c3-measured-feedback-work-packet-v1.md). The follow-on [P2C4 work packet](docs/design/world-intelligence-p2c4-reviewed-impact-disposition-work-packet-v1.md) -freezes the exact proposal-disposition boundary. +freezes the exact proposal-disposition boundary. The stacked +[P2C5 work packet](docs/design/world-intelligence-p2c5-citation-correctness-outcome-work-packet-v1.md) +freezes independently reviewed citation correctness and its citation-preserving negative control. ## Next — trustworthy live orientation diff --git a/docs/audits/world-intelligence-p2c5-citation-correctness-outcome-2026-08-10.md b/docs/audits/world-intelligence-p2c5-citation-correctness-outcome-2026-08-10.md new file mode 100644 index 0000000..befc791 --- /dev/null +++ b/docs/audits/world-intelligence-p2c5-citation-correctness-outcome-2026-08-10.md @@ -0,0 +1,123 @@ +# World Intelligence P2C5 citation-correctness outcome audit — 2026-08-10 + +Status: **stacked candidate evidence only; not a release, SI4 pass, or applied governance change** + +## Source identity + +- World base: P2C4 commit `189c81be1812ee32ffc28148fb63539c66417661` +- World branch: `codex/world-citation-correctness` +- Core base: measured-impact disposition commit + `3c920bb5c411bd9d91a5e2a6c96d4014e9b66763` +- Core dependency: stacked `codex/measured-impact-observed-result` candidate +- Released identity intentionally unchanged: `ace-domain-world-intelligence==0.9.0`, + `ace-core>=0.5.0,<0.6` + +## Exact point-in-time result + +One source-checkout run recorded two independent treatment reviews and two independent control +reviews under product policy +`sha256:ab9b609d9bba9edc4163cdccdbe8761d8494b2bf52789a92f3124b5d3324750c`: + +```text +treatment review 1: citation_correctness_review:27b875cf41e4f417cec5715aebeca73b +treatment review 2: citation_correctness_review:d4c7f54c69289eb258c84d65ad7c996b +control review 1: citation_correctness_review:b1c980ee2df061ab00caa6e3a46d6e7c +control review 2: citation_correctness_review:a2122102390c938dd3edf5104f63f415 +treatment citation coverage: 1.0, 1.0 +control citation coverage: 1.0, 1.0 +treatment correctness: 1.0, 1.0 +control correctness: 0.0, 0.0 +matched pairs: 2 +mean effect: 1.0 +classification: useful +proposal action: promote +proposal live effect: false +historical replay: true +replay reauthorization: false +``` + +The exact evaluation was +`impact_evaluation:eef78b8febd33d4b6121d6f8ce6e1335` with material +`sha256:eef78b8febd33d4b6121d6f8ce6e13354dcb09ef743c5de61940021473595d3a`. +The exact non-effective proposal was +`impact_governance_proposal:d928a69adc1984ff89b1225ed87d93fa` with material +`sha256:d928a69adc1984ff89b1225ed87d93fa0881c3e6f981233d08d9e7756a6c8b80`. + +The negative control preserved both exact citation identities but changed the cited statement from +the admitted publication dates to `2026-15932 published 2026-08-07` and `2026-16197 published +2026-08-06`. It therefore falsifies the structural-coverage proxy without requiring a different +source set. + +## Verification + +Stacked source-checkout verification with the Core observed-result candidate and separately +packaged reference action adapter: + +```text +python -B -m pytest domain_packs/tests/test_p2c3_measured_feedback.py \ + domain_packs/tests/test_p2c4_reviewed_impact_disposition.py \ + domain_packs/tests/test_p2c5_citation_correctness_outcome.py -q --tb=short +6 passed in 1.16s + +ruff check scripts/p2c3_measured_feedback.py \ + scripts/p2c4_reviewed_impact_disposition.py \ + scripts/p2c5_citation_correctness_outcome.py \ + domain_packs/tests/test_p2c5_citation_correctness_outcome.py +PASS + +ruff format --check +PASS + +python -B -m pytest -q --tb=short +89 passed in 15.40s + +python -B -m pytest adapters/federal_register_source/tests -q --tb=short +26 passed in 0.23s + +python -B -m pytest tests/test_release_contract.py -q --tb=short +7 passed in 0.04s + +# Isolated environment with the built World wheel, public ace-core==0.5.0, +# and the separately installed Federal Register source adapter; no Core checkout. +python -B -m pytest -q --tb=short -rs +82 passed, 7 skipped in 15.04s + +ruff check . +PASS + +uv build --out-dir +Successfully built unchanged 0.9.0 source distribution and inert data-only wheel + +git diff --check +PASS +``` + +The seven isolated-environment skips are explicit candidate boundaries: one P2C2 test requires +the separately packaged Core reference action adapter, and two tests each require the unreleased +P2C3, P2C4, and P2C5 Core candidates. The released dependency range and 0.9.0 artifact remain +coherent. Wheel inspection contains only inert Domain Pack JSON and distribution metadata; no +candidate script, test, adapter, or audit Python is shipped. + +Repository-wide `ruff format --check .` reports 19 pre-existing files outside this packet that do +not match the Core candidate environment's formatter version. The four changed Python files pass +the scoped formatting gate and the complete repository passes `ruff check .`; this packet does not +rewrite unrelated World history. + +## Claim boundary + +The World reviewer exact-loads both admitted Observation envelopes and derives the expected +document/date statement from their canonical payloads. Its exact result makes reviewer, source +Observations, policy, claim, citations, verdict, score, time, and limitations inspectable. The Core +Outcome points to that exact result and the evaluator refuses missing or future result provenance +when the criterion requires it. + +This demonstrates deterministic product-policy sensitivity over two recorded official public +records. It does not prove a live request, current freshness, legal truth, reviewer infallibility, +source independence, correction handling, population performance, general Brief quality, causal +impact, or human benefit. The resulting proposal remains non-effective and unapplied. + +## Remaining work + +Contradiction/correction coverage, detection delay, false-alert rate, revision stability, a +materially different Market journey, public artifacts, compatibility/security/release checks, +opt-in live transport, and any explicit proposal application remain future bounded packets. diff --git a/docs/design/world-intelligence-p2c5-citation-correctness-outcome-work-packet-v1.md b/docs/design/world-intelligence-p2c5-citation-correctness-outcome-work-packet-v1.md new file mode 100644 index 0000000..564ccca --- /dev/null +++ b/docs/design/world-intelligence-p2c5-citation-correctness-outcome-work-packet-v1.md @@ -0,0 +1,99 @@ +# World Intelligence P2C5 citation-correctness outcome work packet (v1) + +**Status:** stacked source-checkout candidate; this packet does not release World Intelligence, +apply a governance proposal, close ACE Core issue #38, pass SI4, or complete ACE 0.6.0. + +**Frozen:** 2026-08-10 from World P2C4 commit +`189c81be1812ee32ffc28148fb63539c66417661`, stacked on the Core exact observed-result +provenance candidate. + +## Objective + +Replace P2C3's structural coverage-only outcome with one independently recorded product-quality +result while preserving the full governed journey: + +```text +Observation -> Shift -> Signal -> Brief -> Decision -> reviewed Action + -> exact independent citation review -> observed Outcome + -> useful / harmful / unproven evaluation -> proposal only +``` + +The packet must distinguish presence from correctness. Its negative control retains the exact two +citation identities used by the Brief but swaps the two official publication dates in the cited +claim. Treatment and control therefore both have `1.0` citation coverage while correctness is +`1.0` for the admitted Brief and `0.0` for the semantic-corruption control. + +## Product-owned review policy + +World owns `world_official_record_citation_correctness` version `candidate-1`. The frozen rule +reviews cited claims only, requires the exact two Brief citation identities, exact-loads the two +admitted Observation envelopes, derives the expected document/date statement from their canonical +payloads, and compares the cited statement to those exact recorded facts. Score is supported cited +claims divided by reviewed cited claims. + +The exact review record names: + +- the reviewed Brief or control immutable reference; +- the authenticated reviewer `principal:world-citation-correctness-reviewer`; +- the policy identity, version, and material digest; +- the two exact source Observation references; +- each claim identity, statement, exact citation set, verdict, and rationale; +- coverage, correctness score, limitations, review time, and derived review identity/digest. + +Core and Intelligence see only the result's generic immutable reference from the Outcome measures. +Federal Register, citation, reviewer, and source-policy nouns remain in World. + +## Exact acceptance + +P2C5 must: + +1. rerun P2C2 through P2C4 and preserve the earlier structural result and reject/no-action + disposition as immutable history; +2. append one exact citation-preserving corrupted control artifact; +3. create two distinct reviewed treatment/control Action pairs under matched task conditions; +4. append four independently authenticated review records over the exact subjects and source + Observations; +5. record four Core Outcomes whose measures name the exact review records that produced their + scalar scores; +6. require exact observed-result provenance under a frozen correctness criterion; +7. show treatment/control coverage `1.0/1.0` but correctness `1.0/0.0` in both pairs; +8. classify the bounded result `useful`, emit only a non-effective `promote` proposal, and perform + no proposal application; +9. replay the evaluation without reauthorization; and +10. retain explicit non-claims for network freshness, causality, general Brief quality, human + benefit, and autonomous publication. + +## Negative and failure controls + +The citation-preserving date swap is the primary product negative control: an identifier/string +coverage scorer cannot distinguish it, while the frozen correctness review must. Core's stacked +tests separately require missing observed-result provenance to become unproven, reject cross- +product result coordinates, and exclude post-cutoff result material without loading its payload. +Earlier missing attribution, condition mismatch, Outcome unavailability, duplicate/replay, +interruption, restart, and denied-authority controls remain required. + +## Files and rollback + +This packet owns: + +- `scripts/p2c5_citation_correctness_outcome.py`; +- `domain_packs/tests/test_p2c5_citation_correctness_outcome.py`; +- additive P2C3/P2C4 acceptance-state handoffs; +- this work packet, its audit, and restrained README/roadmap references. + +It changes no shipped Domain Pack, connector, package version, dependency range, lockfile, release +record, or public artifact. Rollback removes the harness, tests, state handoffs, and candidate +documentation. Product review and Outcome records already persisted by a host remain immutable +history. + +## Non-claims and next packet + +The result covers one cited claim, two recorded official records, two matched pairs, and one exact +product rule. It does not validate inference claims, source independence, corrections, live +network freshness, human usefulness, general Brief quality, or causal benefit. The synthetic +control establishes criterion sensitivity, not a population estimate. + +The next bounded outcome packet should measure contradiction/correction coverage or detection +delay over additional public records. Independent Market reproduction, public Core artifacts, +compatibility/security/release gates, opt-in live transport, and any separately authorized proposal +application remain separate work. diff --git a/domain_packs/tests/test_p2c5_citation_correctness_outcome.py b/domain_packs/tests/test_p2c5_citation_correctness_outcome.py new file mode 100644 index 0000000..a82fb2c --- /dev/null +++ b/domain_packs/tests/test_p2c5_citation_correctness_outcome.py @@ -0,0 +1,71 @@ +from __future__ import annotations + +import importlib.util + +import pytest + + +def _require_candidate_contracts() -> None: + if importlib.util.find_spec("ace.application.measured_impact") is None: + pytest.skip("P2C5 requires the stacked ACE Core measured-impact candidate") + from ace.intelligence import ImpactOutcomeMeasuresV1Alpha1 + + if "observed_result" not in ImpactOutcomeMeasuresV1Alpha1.model_fields: + pytest.skip("P2C5 requires exact observed-result provenance from the stacked Core candidate") + if importlib.util.find_spec("ace_reference_workspace_action") is None: + pytest.skip("P2C5 requires the separately packaged Core reference adapter") + + +@pytest.mark.asyncio +async def test_independent_citation_review_becomes_an_exact_measured_outcome(tmp_path) -> None: + _require_candidate_contracts() + from scripts.p2c5_citation_correctness_outcome import run_citation_correctness_outcome + + result = await run_citation_correctness_outcome(tmp_path) + + assert result["review_policy"]["reviewer_ref"] == "principal:world-citation-correctness-reviewer" + assert result["evaluation"]["criterion"]["requires_observed_result"] is True + assert result["evaluation"]["classification"] == "useful" + assert result["evaluation"]["matched_pair_count"] == 2 + assert result["evaluation"]["mean_effect"] == 1.0 + assert result["proposal"]["action"] == "promote" + assert result["proposal"]["live_effect"] is False + assert result["proposal"]["selectable"] is False + assert result["proposal"]["requires_human_review"] is True + assert result["replay"] == { + "historical": True, + "no_reauthorization": True, + "transaction_receipt_id": result["replay"]["transaction_receipt_id"], + } + reviews = (*result["observed_results"]["treatment"], *result["observed_results"]["control"]) + assert {item["contract"] for item in reviews} == {"ace.world-intelligence.citation-correctness-review/v1alpha1"} + assert {item["reviewer_context"]["actor_ref"] for item in reviews} == { + "principal:world-citation-correctness-reviewer" + } + assert {len(item["source_observations"]) for item in reviews} == {2} + + +@pytest.mark.asyncio +async def test_citation_preserving_negative_control_separates_coverage_from_correctness(tmp_path) -> None: + _require_candidate_contracts() + from scripts.p2c5_citation_correctness_outcome import run_citation_correctness_outcome + + result = await run_citation_correctness_outcome(tmp_path) + control = result["negative_control"] + + assert len(control["citation_ids_preserved"]) == 2 + assert control["treatment_citation_coverage"] == (1.0, 1.0) + assert control["control_citation_coverage"] == (1.0, 1.0) + assert control["treatment_correctness"] == (1.0, 1.0) + assert control["control_correctness"] == (0.0, 0.0) + assert "2026-15932 published 2026-08-07" in control["corrupted_statement"] + assert result["scope"] == { + "independent_exact_review": True, + "recorded_official_sources": True, + "network_freshness_claimed": False, + "general_brief_quality_claimed": False, + "human_benefit_claimed": False, + "causality_claimed": False, + "proposal_applied": False, + "autonomous_publication": False, + } diff --git a/scripts/p2c3_measured_feedback.py b/scripts/p2c3_measured_feedback.py index 6890b95..1fc3d9f 100644 --- a/scripts/p2c3_measured_feedback.py +++ b/scripts/p2c3_measured_feedback.py @@ -753,6 +753,10 @@ async def run_measured_feedback( state_sink.update(state) state_sink.update( { + "impact_target_ref": target_ref, + "impact_control_ref": control_ref, + "measured_treatments": treatments, + "measured_controls": controls, "impact_binding": impact_binding, "impact_criterion": criterion, "impact_request": request, diff --git a/scripts/p2c4_reviewed_impact_disposition.py b/scripts/p2c4_reviewed_impact_disposition.py index 6d28b5b..3ef554f 100644 --- a/scripts/p2c4_reviewed_impact_disposition.py +++ b/scripts/p2c4_reviewed_impact_disposition.py @@ -74,7 +74,11 @@ def _install_disposition_policy(state: dict[str, Any]) -> GovernedOperationBindi return binding -async def run_reviewed_disposition(workspace_root: Path) -> dict[str, Any]: +async def run_reviewed_disposition( + workspace_root: Path, + *, + state_sink: dict[str, Any] | None = None, +) -> dict[str, Any]: """Run P2C3, then record one exact reject/no-action human Decision.""" state: dict[str, Any] = {} @@ -125,7 +129,7 @@ async def run_reviewed_disposition(workspace_root: Path) -> dict[str, Any]: if heads_after != heads_before: raise AssertionError("reviewed proposal disposition mutated effective governed state") - return { + result = { "contract": "ace.world-intelligence.reviewed-impact-disposition/v1alpha1", "measured_feedback": measured, "disposition": { @@ -148,6 +152,16 @@ async def run_reviewed_disposition(workspace_root: Path) -> dict[str, Any]: "autonomous_publication": False, }, } + if state_sink is not None: + state_sink.update(state) + state_sink.update( + { + "impact_disposition_binding": binding, + "impact_disposition_request": request, + "impact_disposition_admission": admission, + } + ) + return result def main() -> None: diff --git a/scripts/p2c5_citation_correctness_outcome.py b/scripts/p2c5_citation_correctness_outcome.py new file mode 100644 index 0000000..88b1e1a --- /dev/null +++ b/scripts/p2c5_citation_correctness_outcome.py @@ -0,0 +1,705 @@ +"""Independently review exact citation correctness and measure its impact.""" + +from __future__ import annotations + +import asyncio +import json +from datetime import datetime, timedelta +from pathlib import Path +from typing import Any, Literal, Self + +from pydantic import BaseModel, ConfigDict, Field, field_validator, model_validator + +from ace.application import MeasuredImpactService +from ace.core import ( + AppendOnlyTransactionRequestV1, + AuthenticatedRuntimeContextV1Alpha1, + CapabilityArtifactIdentityV1Alpha1, + GovernedOperationBindingV1Alpha1, + GovernedStateHeadPreconditionV1Alpha1, + ImmutableRecordReferenceV1, + ImmutableRecordV1, + OutcomeIntentV1Alpha1, + OutcomeV1Alpha1, + canonical_hash, + canonical_json, + capability_state_ref_for_artifact, +) +from ace.intelligence import ( + ImpactClassification, + ImpactConditionsV1Alpha1, + ImpactCriterionV1Alpha1, + ImpactEvaluationRequestV1Alpha1, + ImpactEvidenceV1Alpha1, + ImpactGovernanceAction, + ImpactMetricDirection, + ImpactOutcomeMeasuresV1Alpha1, + ImpactTargetKind, +) + +from scripts.p2c2_federal_register_monitor import _time +from scripts.p2c2_governed_reality_brief import _context, _head +from scripts.p2c3_measured_feedback import ( + ReviewedExport, + _append_value, + _authorize_append, + _digest, + _record_attribution, + _run_reviewed_export, +) +from scripts.p2c4_reviewed_impact_disposition import run_reviewed_disposition + +OUTCOME_TYPE = "independent_artifact_review" +MEASURE_ID = "official_citation_correctness" +CRITERION_ID = "impact_criterion:world-official-citation-correctness" +CRITERION_FROZEN_AT = _time("2026-08-07T18:01:08Z") +REVIEW_POLICY_ID = "world_official_record_citation_correctness" +REVIEW_POLICY_VERSION = "candidate-1" + +IMPACT_ARTIFACT = CapabilityArtifactIdentityV1Alpha1( + capability="measured_impact_evaluation", + contract="ace.application.measured-impact-service/v1alpha1", + implementation_id="world_citation_correctness_candidate", + implementation_version="0.1.0", + artifact_digest="sha256:" + "e" * 64, +) + + +class _FrozenModel(BaseModel): + model_config = ConfigDict(extra="forbid", frozen=True) + + +def _derive_identity(value: _FrozenModel, *, prefix: str, id_field: str, digest_field: str) -> None: + material = value.model_dump(mode="json", exclude={id_field, digest_field}) + digest = canonical_hash(material) + expected_id = f"{prefix}:{digest[:32]}" + expected_digest = f"sha256:{digest}" + supplied_id = getattr(value, id_field) + supplied_digest = getattr(value, digest_field) + if supplied_id is not None and supplied_id != expected_id: + raise ValueError(f"{id_field} does not match exact review material") + if supplied_digest is not None and supplied_digest != expected_digest: + raise ValueError(f"{digest_field} does not match exact review material") + object.__setattr__(value, id_field, expected_id) + object.__setattr__(value, digest_field, expected_digest) + + +class CitationClaimAssessmentV1Alpha1(_FrozenModel): + """Product-owned correctness judgment for one exact cited claim.""" + + contract: Literal["ace.world-intelligence.citation-claim-assessment/v1alpha1"] = ( + "ace.world-intelligence.citation-claim-assessment/v1alpha1" + ) + claim_id: str + statement: str + citation_ids: tuple[str, ...] = Field(min_length=1, max_length=16) + verdict: Literal["supported", "unsupported"] + rationale: str + assessment_id: str | None = None + assessment_digest: str | None = None + + @field_validator("citation_ids") + @classmethod + def canonicalize_citations(cls, value: tuple[str, ...]) -> tuple[str, ...]: + ordered = tuple(sorted(value)) + if len(ordered) != len(set(ordered)): + raise ValueError("citation assessment cannot amplify duplicate citation identities") + return ordered + + @model_validator(mode="after") + def derive_identity(self) -> Self: + _derive_identity( + self, + prefix="citation_claim_assessment", + id_field="assessment_id", + digest_field="assessment_digest", + ) + return self + + +class CitationCorrectnessReviewV1Alpha1(_FrozenModel): + """Exact independent review result later named by a Core Outcome.""" + + contract: Literal["ace.world-intelligence.citation-correctness-review/v1alpha1"] = ( + "ace.world-intelligence.citation-correctness-review/v1alpha1" + ) + product_id: str + review_key: str + reviewed_subject: ImmutableRecordReferenceV1 + reviewer_context: AuthenticatedRuntimeContextV1Alpha1 + policy_id: str + policy_version: str + policy_digest: str + source_observations: tuple[ImmutableRecordReferenceV1, ...] = Field(min_length=1, max_length=16) + assessments: tuple[CitationClaimAssessmentV1Alpha1, ...] = Field(min_length=1, max_length=16) + citation_coverage: float = Field(ge=0.0, le=1.0) + correctness_score: float = Field(ge=0.0, le=1.0) + limitations: tuple[str, ...] + reviewed_at: datetime + review_id: str | None = None + review_digest: str | None = None + + @field_validator("source_observations") + @classmethod + def canonicalize_observations( + cls, value: tuple[ImmutableRecordReferenceV1, ...] + ) -> tuple[ImmutableRecordReferenceV1, ...]: + ordered = tuple(sorted(value, key=lambda item: item.storage_id)) + identities = tuple(item.storage_id for item in ordered) + if len(identities) != len(set(identities)): + raise ValueError("citation review cannot amplify duplicate Observation identities") + return ordered + + @model_validator(mode="after") + def validate_scope_scores_and_identity(self) -> Self: + if ( + self.reviewed_subject.product_id != self.product_id + or self.reviewer_context.product_id != self.product_id + or any(item.product_id != self.product_id for item in self.source_observations) + ): + raise ValueError("citation correctness review crossed exact product scope") + if len({item.assessment_id for item in self.assessments}) != len(self.assessments): + raise ValueError("citation correctness review duplicated an assessment identity") + expected_score = sum(item.verdict == "supported" for item in self.assessments) / len(self.assessments) + if self.correctness_score != expected_score: + raise ValueError("citation correctness score differs from exact assessments") + _derive_identity( + self, + prefix="citation_correctness_review", + id_field="review_id", + digest_field="review_digest", + ) + return self + + +class _ReplayMustNotAuthorize: + async def authorize_action(self, request): + raise AssertionError(f"historical citation evaluation requested new authority: {request.authorization_key}") + + +def _install_policy(state: dict[str, Any]): + environment = state["environment"] + runtime = state["runtime"] + product_id = environment.fixture["product_id"] + criterion_head = _head(product_id, "impact_criterion", CRITERION_ID, 60) + operation_head = _head( + product_id, + "governed_operation_configuration", + "governed_operation_configuration:world-citation-correctness", + 61, + ) + binding = GovernedOperationBindingV1Alpha1( + product_id=product_id, + artifact=IMPACT_ARTIFACT, + configuration_ref=operation_head.state_id, + authority="append_measured_impact", + grant_ref="authority_grant:world-citation-correctness", + state_head_precondition=GovernedStateHeadPreconditionV1Alpha1.from_head(operation_head), + ) + capability_head = _head(product_id, "capability_state", capability_state_ref_for_artifact(IMPACT_ARTIFACT), 62) + authority_head = _head(product_id, "authority_grant", binding.grant_ref, 63) + for head in (criterion_head, operation_head, capability_head, authority_head): + environment.store.set_governed_state_head(head) + runtime.heads[head.state_kind, head.state_id] = head + runtime.bindings = (*runtime.bindings, binding) + return criterion_head, binding + + +async def _append_corrupted_control( + state: dict[str, Any], + *, + citation_ids: tuple[str, ...], + observation_refs: tuple[ImmutableRecordReferenceV1, ...], + corrupted_statement: str, +) -> tuple[ImmutableRecordReferenceV1, str, tuple[dict[str, Any], ...]]: + environment = state["environment"] + content = ( + "# Citation-Preserving Correctness Control\n\n" + f"{corrupted_statement}\n\n" + f"Exact citation identities retained: {', '.join(citation_ids)}.\n\n" + "Exact admitted Observation identities retained: " + f"{', '.join(item.record_key for item in observation_refs)}.\n" + ) + claims = ( + { + "claim_id": f"corrupted_claim:{canonical_hash([corrupted_statement, citation_ids])[:32]}", + "grounding_kind": "cited", + "statement": corrupted_statement, + "citation_ids": citation_ids, + }, + ) + payload = { + "control_type": "citation_preserving_semantic_corruption", + "content_markdown": content, + "claims": claims, + "citation_ids": citation_ids, + "observation_references": [item.model_dump(mode="json") for item in observation_refs], + "limitations": ["synthetic_negative_control_over_exact_recorded_public_sources"], + } + requested_at = state["clock"]() + authorization = await _authorize_append( + state, + context=environment.context, + authorization_key="append:world-citation-preserving-control", + subject_ref="world_citation_preserving_control:2026-08-07", + subject_digest=_digest(payload), + requested_at=requested_at, + ) + record = ImmutableRecordV1( + product_id=environment.fixture["product_id"], + record_space="world_intelligence", + record_kind="brief_control", + record_key="brief_control:citation-preserving-corruption:2026-08-07", + payload_contract="ace.world-intelligence.citation-preserving-control/v1alpha1", + payload=payload, + as_of=state["impact_target_ref"].as_of, + available_at=authorization.authorized_at, + processing_order=0, + ) + append = AppendOnlyTransactionRequestV1( + product_id=record.product_id, + record_space=record.record_space, + transaction_key="world-citation-preserving-control:2026-08-07", + records=(record,), + submitted_at=record.available_at, + governed_state_preconditions=authorization.state_preconditions, + ) + if await environment.store.append(append) != append.receipt(): + raise AssertionError("citation-preserving control append returned divergent material") + return record.reference(), content, claims + + +def _policy_digest( + *, + expected_statement: str, + citation_ids: tuple[str, ...], + observation_refs: tuple[ImmutableRecordReferenceV1, ...], +) -> str: + return _digest( + { + "policy_id": REVIEW_POLICY_ID, + "policy_version": REVIEW_POLICY_VERSION, + "expected_statement": expected_statement, + "expected_citation_ids": citation_ids, + "source_observations": [item.model_dump(mode="json") for item in observation_refs], + "score": "supported_cited_claims / reviewed_cited_claims", + } + ) + + +async def _statements_from_observations( + state: dict[str, Any], observation_refs: tuple[ImmutableRecordReferenceV1, ...] +) -> tuple[str, str]: + """Derive the positive and date-swapped claims from exact admitted records.""" + + facts: list[tuple[str, str]] = [] + for reference in observation_refs: + record = await state["environment"].store.load_record( + reference.storage_id, + product_id=reference.product_id, + record_space=reference.record_space, + record_kind=reference.record_kind, + ) + if ( + record is None + or record.reference() != reference + or record.payload_contract != "ace.intelligence.observation/v1alpha1" + ): + raise AssertionError("citation review source Observation is unavailable or changed") + try: + value = json.loads(record.payload["payload"]["value_json"]) + facts.append((value["document_number"], value["publication_date"])) + except (KeyError, TypeError, json.JSONDecodeError): + raise AssertionError("citation review Observation lost required official record facts") from None + facts.sort() + if len(facts) != 2 or len({number for number, _ in facts}) != 2: + raise AssertionError("citation review requires two distinct exact official records") + first, second = facts + expected = ( + f"The admitted official records are {first[0]} published {first[1]} and {second[0]} published {second[1]}." + ) + corrupted = ( + f"The admitted official records are {first[0]} published {second[1]} and {second[0]} published {first[1]}." + ) + return expected, corrupted + + +async def _review_artifact( + state: dict[str, Any], + *, + subject: ImmutableRecordReferenceV1, + claims: tuple[dict[str, Any], ...], + citation_ids: tuple[str, ...], + observation_refs: tuple[ImmutableRecordReferenceV1, ...], + expected_statement: str, + pair_index: int, + variant: str, +) -> tuple[CitationCorrectnessReviewV1Alpha1, ImmutableRecordReferenceV1]: + environment = state["environment"] + cited_claims = tuple(item for item in claims if item["grounding_kind"] == "cited") + if not cited_claims: + raise AssertionError("citation correctness review requires an exact cited claim") + expected_ids = tuple(sorted(citation_ids)) + assessments = tuple( + CitationClaimAssessmentV1Alpha1( + claim_id=item["claim_id"], + statement=item["statement"], + citation_ids=tuple(item["citation_ids"]), + verdict=( + "supported" + if item["statement"] == expected_statement and tuple(sorted(item["citation_ids"])) == expected_ids + else "unsupported" + ), + rationale=( + "The exact cited statement matches the product-frozen facts and exact citation set." + if item["statement"] == expected_statement and tuple(sorted(item["citation_ids"])) == expected_ids + else "Citation identities are present, but the statement contradicts the product-frozen publication dates." + ), + ) + for item in cited_claims + ) + present_ids = {citation_id for item in cited_claims for citation_id in item["citation_ids"]} + coverage = len(present_ids & set(expected_ids)) / len(expected_ids) + reviewed_at = state["clock"]() + reviewer = _context(environment.context, "principal:world-citation-correctness-reviewer") + review = CitationCorrectnessReviewV1Alpha1( + product_id=environment.fixture["product_id"], + review_key=f"citation-correctness-review:{pair_index}:{variant}", + reviewed_subject=subject, + reviewer_context=reviewer, + policy_id=REVIEW_POLICY_ID, + policy_version=REVIEW_POLICY_VERSION, + policy_digest=_policy_digest( + expected_statement=expected_statement, + citation_ids=expected_ids, + observation_refs=observation_refs, + ), + source_observations=observation_refs, + assessments=assessments, + citation_coverage=coverage, + correctness_score=sum(item.verdict == "supported" for item in assessments) / len(assessments), + limitations=( + "bounded_to_one_exact_cited_claim_and_two_recorded_official_sources", + "does_not_establish_general_brief_quality_or_human_benefit", + ), + reviewed_at=reviewed_at, + ) + authorization = await _authorize_append( + state, + context=reviewer, + authorization_key=f"citation-correctness-review:{pair_index}:{variant}", + subject_ref=str(review.review_id), + subject_digest=str(review.review_digest), + requested_at=reviewed_at, + ) + reference = await _append_value( + state, + value=review, + record_kind="citation_correctness_review", + record_key=str(review.review_id), + transaction_key=f"citation-correctness-review:{review.review_id}", + as_of=reviewed_at, + authorization=authorization, + ) + return review, reference + + +async def _record_review_outcome( + state: dict[str, Any], + *, + export: ReviewedExport, + review: CitationCorrectnessReviewV1Alpha1, + review_ref: ImmutableRecordReferenceV1, + pair_index: int, + variant: str, +) -> ImmutableRecordReferenceV1: + environment = state["environment"] + measures = ImpactOutcomeMeasuresV1Alpha1( + primary_value=review.correctness_score, + observed_result=review_ref, + latency_ms=max(0, int((review.reviewed_at - export.intent.requested_at).total_seconds() * 1_000)), + cost_usd=0.0, + failure_count=0, + degraded=False, + limitations=review.limitations, + ) + recorded_at = state["clock"]() + observer = _context(environment.context, "principal:world-citation-outcome-observer") + intent = OutcomeIntentV1Alpha1( + product_id=environment.fixture["product_id"], + authenticated_context=observer, + decision=export.decision_ref, + outcome_type=OUTCOME_TYPE, + measure_id=MEASURE_ID, + value_json=canonical_json(measures.model_dump(mode="json")), + observed_at=review.reviewed_at, + recorded_at=recorded_at, + ) + authorization = await _authorize_append( + state, + context=observer, + authorization_key=f"citation-correctness-outcome:{pair_index}:{variant}", + subject_ref=str(intent.intent_id), + subject_digest=str(intent.intent_digest), + requested_at=recorded_at, + ) + outcome = OutcomeV1Alpha1(intent=intent, authorization=authorization) + return await _append_value( + state, + value=outcome, + record_kind="outcome", + record_key=str(outcome.outcome_id), + transaction_key=f"outcome:{outcome.outcome_id}", + as_of=review.reviewed_at, + authorization=authorization, + ) + + +async def run_citation_correctness_outcome(workspace_root: Path) -> dict[str, Any]: + """Run the P2C4 journey through exact independent citation review Outcomes.""" + + state: dict[str, Any] = {} + prior = await run_reviewed_disposition(workspace_root, state_sink=state) + environment = state["environment"] + target_ref = state["impact_target_ref"] + treatment_exports = tuple(item.export for item in state["measured_treatments"]) + brief = state["brief_admission"].brief + observation_refs = tuple( + record + for admission in state["admissions"] + for record in admission.transaction_receipt.records + if record.record_kind == "observation" + ) + citation_ids = tuple(sorted(item.citation_id for item in brief.citations)) + treatment_claims = tuple(item.model_dump(mode="json") for item in brief.claims) + expected_statement, corrupted_statement = await _statements_from_observations(state, observation_refs) + control_ref, control_content, control_claims = await _append_corrupted_control( + state, + citation_ids=citation_ids, + observation_refs=observation_refs, + corrupted_statement=corrupted_statement, + ) + control_exports = tuple( + [ + await _run_reviewed_export( + state, + subject=control_ref, + content=control_content, + pair_index=index, + variant="citation-control", + ) + for index in (1, 2) + ] + ) + criterion_head, impact_binding = _install_policy(state) + criterion = ImpactCriterionV1Alpha1( + product_id=environment.fixture["product_id"], + criterion_id=CRITERION_ID, + criterion_version="candidate-1", + target_kind=ImpactTargetKind.INTELLIGENCE_ARTIFACT, + outcome_type=OUTCOME_TYPE, + measure_id=MEASURE_ID, + metric_direction=ImpactMetricDirection.HIGHER_IS_BETTER, + useful_effect_threshold=0.5, + harmful_effect_threshold=0.5, + minimum_matched_pairs=2, + requires_observed_result=True, + harmful_action=ImpactGovernanceAction.ROLLBACK, + state_head_precondition=GovernedStateHeadPreconditionV1Alpha1.from_head(criterion_head), + frozen_at=CRITERION_FROZEN_AT, + ) + + evidence: list[ImpactEvidenceV1Alpha1] = [] + treatment_reviews: list[CitationCorrectnessReviewV1Alpha1] = [] + control_reviews: list[CitationCorrectnessReviewV1Alpha1] = [] + observed_times: list[datetime] = [] + for index, (treatment_export, control_export) in enumerate( + zip(treatment_exports, control_exports, strict=True), start=1 + ): + treatment_attribution = await _record_attribution( + state, + export=treatment_export, + subject=target_ref, + pair_index=index, + variant="citation-treatment", + ) + control_attribution = await _record_attribution( + state, + export=control_export, + subject=control_ref, + pair_index=index, + variant="citation-control", + ) + treatment_review, treatment_review_ref = await _review_artifact( + state, + subject=target_ref, + claims=treatment_claims, + citation_ids=citation_ids, + observation_refs=observation_refs, + expected_statement=expected_statement, + pair_index=index, + variant="treatment", + ) + control_review, control_review_ref = await _review_artifact( + state, + subject=control_ref, + claims=control_claims, + citation_ids=citation_ids, + observation_refs=observation_refs, + expected_statement=expected_statement, + pair_index=index, + variant="control", + ) + treatment_outcome = await _record_review_outcome( + state, + export=treatment_export, + review=treatment_review, + review_ref=treatment_review_ref, + pair_index=index, + variant="treatment", + ) + control_outcome = await _record_review_outcome( + state, + export=control_export, + review=control_review, + review_ref=control_review_ref, + pair_index=index, + variant="control", + ) + treatment_reviews.append(treatment_review) + control_reviews.append(control_review) + observed_times.extend((treatment_review.reviewed_at, control_review.reviewed_at)) + conditions = ImpactConditionsV1Alpha1( + product_id=environment.fixture["product_id"], + condition_key=f"world-citation-correctness-pair:{index}", + route_id="world:fcc-independent-citation-review", + context_json=canonical_json( + { + "citation_count": len(citation_ids), + "pair_index": index, + "recorded_transport": True, + "review_policy_digest": treatment_review.policy_digest, + "task": "independent_exact_citation_correctness_review", + } + ), + observation_window_start=CRITERION_FROZEN_AT + timedelta(seconds=1), + observation_window_end=max(observed_times) + timedelta(minutes=1), + frozen_at=CRITERION_FROZEN_AT, + ) + evidence.append( + ImpactEvidenceV1Alpha1( + product_id=environment.fixture["product_id"], + evidence_key=f"world-citation-correctness:{index}", + treatment_attribution=treatment_attribution, + control_attribution=control_attribution, + treatment_decision=treatment_export.decision_ref, + control_decision=control_export.decision_ref, + treatment_action_review=treatment_export.review_ref, + treatment_action_admission=treatment_export.admission_ref, + treatment_action_terminal=treatment_export.terminal_ref, + control_action_review=control_export.review_ref, + control_action_admission=control_export.admission_ref, + control_action_terminal=control_export.terminal_ref, + treatment_outcome=treatment_outcome, + control_outcome=control_outcome, + treatment_conditions=conditions, + control_conditions=conditions, + ) + ) + + cutoff_at = state["clock"]() + request = ImpactEvaluationRequestV1Alpha1( + evaluation_key="world-citation-correctness:fcc-publication-change:2026-08-07", + product_id=environment.fixture["product_id"], + authenticated_context=environment.context, + criterion=criterion, + target=target_ref, + control=control_ref, + evidence=tuple(evidence), + cutoff_at=cutoff_at, + requested_at=state["clock"](), + ) + service = MeasuredImpactService( + store=environment.store, + authorizer=state["reasoning"], + operation_binding=impact_binding, + ) + admission = await service.evaluate(request) + replay = await MeasuredImpactService( + store=environment.store, + authorizer=_ReplayMustNotAuthorize(), + operation_binding=impact_binding, + ).evaluate(request) + if admission.replayed or not replay.replayed or replay.evaluation != admission.evaluation: + raise AssertionError("citation correctness evaluation did not replay exact historical material") + if admission.evaluation.classification is not ImpactClassification.USEFUL: + raise AssertionError("frozen citation correctness criterion did not classify useful") + if admission.proposal is None or admission.proposal.action is not ImpactGovernanceAction.PROMOTE: + raise AssertionError("citation correctness result did not emit its proposal-only mapping") + if any(item.citation_coverage != 1.0 for item in (*treatment_reviews, *control_reviews)): + raise AssertionError("citation-preserving control lost an exact citation identity") + if {item.correctness_score for item in treatment_reviews} != {1.0}: + raise AssertionError("supported treatment citation did not score correct") + if {item.correctness_score for item in control_reviews} != {0.0}: + raise AssertionError("semantic-corruption control did not score incorrect") + + return { + "contract": "ace.world-intelligence.citation-correctness-outcome/v1alpha1", + "prior_reviewed_disposition": prior, + "review_policy": { + "policy_id": REVIEW_POLICY_ID, + "policy_version": REVIEW_POLICY_VERSION, + "policy_digest": treatment_reviews[0].policy_digest, + "reviewer_ref": treatment_reviews[0].reviewer_context.actor_ref, + "expected_statement": expected_statement, + }, + "negative_control": { + "control_reference": control_ref.model_dump(mode="json"), + "corrupted_statement": corrupted_statement, + "citation_ids_preserved": citation_ids, + "treatment_citation_coverage": tuple(item.citation_coverage for item in treatment_reviews), + "control_citation_coverage": tuple(item.citation_coverage for item in control_reviews), + "treatment_correctness": tuple(item.correctness_score for item in treatment_reviews), + "control_correctness": tuple(item.correctness_score for item in control_reviews), + }, + "observed_results": { + "treatment": tuple(item.model_dump(mode="json") for item in treatment_reviews), + "control": tuple(item.model_dump(mode="json") for item in control_reviews), + }, + "evaluation": admission.evaluation.model_dump(mode="json"), + "proposal": admission.proposal.model_dump(mode="json"), + "replay": { + "historical": replay.replayed, + "no_reauthorization": True, + "transaction_receipt_id": str(replay.transaction_receipt.receipt_id), + }, + "scope": { + "independent_exact_review": True, + "recorded_official_sources": True, + "network_freshness_claimed": False, + "general_brief_quality_claimed": False, + "human_benefit_claimed": False, + "causality_claimed": False, + "proposal_applied": False, + "autonomous_publication": False, + }, + } + + +def main() -> None: + import argparse + + parser = argparse.ArgumentParser() + parser.add_argument("workspace_root", type=Path) + args = parser.parse_args() + print( + json.dumps( + asyncio.run(run_citation_correctness_outcome(args.workspace_root)), + indent=2, + sort_keys=True, + ) + ) + + +if __name__ == "__main__": + main()