diff --git a/README.md b/README.md index d40a6c3..ae000a7 100644 --- a/README.md +++ b/README.md @@ -248,6 +248,19 @@ of `1.0` and `0.0` across two matched pairs and another non-effective proposal. This freezes a World product rule over one recorded pair. It is not live revision, a general semantic-equivalence engine, population stability, calibration, causality, or human benefit. +P2C9 freezes one exact probabilistic scoring boundary over that recorded correction. After the +original Observation is admitted but before the correction is available, treatment and control +record probabilities of `0.75` and `0.25` for the same explicit-correction event and complete the +same reviewed workflow. Only then is the exact correction Observation admitted. Independent +reviews derive the binary result and single-event Brier quality contributions of `0.9375` and +`0.4375`; the two-pair difference is useful and still emits only a non-effective proposal. + +The forecast artifacts contain no correction identity or result material, and every reviewed +forecast Action completes before the result becomes available. The probabilities are declared +fixture inputs, not ACE outputs. This is exact forecast/result scoring, not a historically +contemporaneous forecast, model-skill finding, empirical calibration curve, population reliability +estimate, causal claim, or human-benefit finding. + ## What the public World proof demonstrates Generate a self-contained visual Reality Brief and its exact machine-readable backing data: @@ -342,11 +355,12 @@ $PY -m scripts.p2c3_measured_feedback "$WORKSPACE" # Stacked candidate: explicit reject/no-action review of the exact proposal $PY -m scripts.p2c4_reviewed_impact_disposition "$WORKSPACE" -# Stacked candidates: correctness, attention, correction delay, and revision stability +# Stacked candidates: correctness, attention, correction delay, revision stability, and forecast scoring $PY -m scripts.p2c5_citation_correctness_outcome "$WORKSPACE" $PY -m scripts.p2c6_contradiction_attention_outcome "$WORKSPACE" $PY -m scripts.p2c7_correction_detection_delay_outcome "$WORKSPACE" $PY -m scripts.p2c8_correction_revision_stability_outcome "$WORKSPACE" +$PY -m scripts.p2c9_forecast_calibration_outcome "$WORKSPACE" ``` The released 0.9.0 gates are reproducible through the locked environment, as CI does. The @@ -383,6 +397,8 @@ The complete governed product-journey evidence is recorded in [`docs/audits/world-intelligence-p2c2-governed-reality-brief-2026-08-10.md`](docs/audits/world-intelligence-p2c2-governed-reality-brief-2026-08-10.md). The source-checkout measured-feedback candidate is recorded in [`docs/audits/world-intelligence-p2c3-measured-feedback-2026-08-10.md`](docs/audits/world-intelligence-p2c3-measured-feedback-2026-08-10.md). +The latest withheld-result forecast-scoring candidate is recorded in +[`docs/audits/world-intelligence-p2c9-forecast-calibration-outcome-2026-08-10.md`](docs/audits/world-intelligence-p2c9-forecast-calibration-outcome-2026-08-10.md). Release-level scope and evidence are recorded in [`docs/releases/world-intelligence-p2c2-v0.9.0.md`](docs/releases/world-intelligence-p2c2-v0.9.0.md), @@ -417,8 +433,10 @@ control, and valid silence under one frozen recorded-source challenge. P2C7 adds one explicit recorded correction pair, preserves the prior record, and measures frozen-replay detection delay against a product target. P2C8 measures exact unaffected-claim identity preservation across real Brief contracts while holding correction semantics, source coverage, and -claim count constant. The next bounded measurement work is calibration, another independently -sourced correction event, or independent Market reproduction. +claim count constant. P2C9 adds an exact withheld-result forecast record and derives a single-event +Brier contribution while explicitly withholding any population-calibration or model-skill claim. +The next bounded measurement work is another independently sourced correction event or independent +Market reproduction. Separately reviewed opt-in network transport and P2D multi-source conflict/correction with LIVE inputs remain independent work. None of these steps may add autonomous publishing, delivery, persuasion, or action authority to a Domain Pack. diff --git a/ROADMAP.md b/ROADMAP.md index 70ca3f4..673fb8c 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -16,7 +16,7 @@ source code into the platform. See the [0.9.0 release record](docs/releases/world-intelligence-p2c2-v0.9.0.md). -## Candidate — P2C3–P2C8 measured feedback and product-owned outcomes +## Candidate — P2C3–P2C9 measured feedback and product-owned outcomes - The exact 0.9.0 Brief and a World-owned source-only control pass through two matched reviewed export pairs under one frozen structural citation-coverage criterion. @@ -46,6 +46,11 @@ See the [0.9.0 release record](docs/releases/world-intelligence-p2c2-v0.9.0.md). workflow, but treatment preserves both unaffected claim identities while the paraphrase-drift control preserves neither. Exact reviews expose the affected, replacement, stable, preserved, drifted, and unexpected claim sets without claiming general semantic equivalence. +- A sixth frozen criterion records treatment and control probabilities before the correction result + is available and completes every reviewed forecast Action first. The later exact correction + resolves the binary event; exact reviews derive single-event Brier quality of `0.9375` and + `0.4375`. This proves forecast/result scoring and leakage-resistant record order, not historical + contemporaneity, probability generation by ACE, model skill, or population calibration. - This is source-checkout evidence against stacked Core candidates, not a released World capability, human benefit finding, causal claim, network-freshness proof, or applied governance change. @@ -63,7 +68,10 @@ freezes exact correction linkage, prior-record preservation, and recorded-replay without claiming live monitoring or network-arrival latency. The stacked [P2C8 work packet](docs/design/world-intelligence-p2c8-correction-revision-stability-outcome-work-packet-v1.md) freezes correction-induced Brief revision stability without claiming live revision, population -performance, or a general semantic-equivalence engine. +performance, or a general semantic-equivalence engine. The stacked +[P2C9 work packet](docs/design/world-intelligence-p2c9-forecast-calibration-outcome-work-packet-v1.md) +freezes an exact forecast/result scoring boundary without claiming a historical forecast, model +skill, or population calibration. ## Next — trustworthy live orientation diff --git a/docs/audits/world-intelligence-p2c9-forecast-calibration-outcome-2026-08-10.md b/docs/audits/world-intelligence-p2c9-forecast-calibration-outcome-2026-08-10.md new file mode 100644 index 0000000..f617754 --- /dev/null +++ b/docs/audits/world-intelligence-p2c9-forecast-calibration-outcome-2026-08-10.md @@ -0,0 +1,155 @@ +# World Intelligence P2C9 forecast-calibration outcome audit — 2026-08-10 + +Status: **stacked candidate evidence only; not a release, SI4 pass, population-calibration proof, +model-skill finding, or applied governance change** + +## Source identity + +- World base: P2C8 commit `129767d27d4af22dba292deb9e691375a0695bb8` +- World branch: `codex/world-forecast-calibration` +- Core dependency: exact observed-result candidate commit + `433e3d16c5458c975557dcd1552824fb959d4d12` +- Released identity intentionally unchanged: `ace-domain-world-intelligence==0.9.0`, + `ace-core>=0.5.0,<0.6` + +## Exact point-in-time result + +The source fixture remains +`sha256:2b81d3950cbfd127408eec227ec5cd249677a189120d6ca7b603d85d01074543`, +with original Observation `observation:f1768d6f4191a86e245846a9a1e33768` and correction +Observation `observation:fced5d3bbc3802c0285021142b332e29`. + +One source-checkout run recorded this exact availability order: + +```text +original Observation available: 2026-08-11T03:15:24.023307Z +forecast issued: 2026-08-11T03:15:24.024601Z +latest reviewed Action done: 2026-08-11T03:15:24.057528Z +correction result available: 2026-08-11T03:15:24.061041Z +``` + +The two forecast records were: + +```text +treatment: + public_event_forecast:cadabc0db05387ffda21e167a7ae8a0c + sha256:cadabc0db05387ffda21e167a7ae8a0cfba830841d39b9141f4faa1cfca4f56c + probability: 0.75 +control: + public_event_forecast:2ebf7bd7c21363bcc6074bd654f020de + sha256:2ebf7bd7c21363bcc6074bd654f020de32a021280450cde5f889e92f24243488 + probability: 0.25 +``` + +Both records bind only the original Observation, the same target event, policy, source family, +resolution rule, and window. Neither contains correction document `2021-10670`, the correction +Observation coordinate, the eventual outcome, or an outcome score. The correction's exact immutable +reference becomes available after all four forecast Actions complete. + +The reviews used World policy `world_recorded_binary_forecast_brier_quality` version `candidate-1`, +material `sha256:759c13b02bff4b6d749ff20888a9fa7d4f3a67e6ef94439801d28d9d99abf4e7`: + +```text +treatment review 1: forecast_resolution_review:dceee9319c93d8f4aaa1712f6ad4a736 +treatment review 2: forecast_resolution_review:388863533f4bceb49c49326c029048f9 +control review 1: forecast_resolution_review:585c2385de9b45b3aeee5705a577cfb0 +control review 2: forecast_resolution_review:6afe8778c77d840aca6b51d5a7198c3e + +binary event outcome: 1.0 +treatment Brier loss: 0.0625, 0.0625 +treatment quality: 0.9375, 0.9375 +control Brier loss: 0.5625, 0.5625 +control quality: 0.4375, 0.4375 +matched pairs: 2 +mean effect: 0.5 +classification: useful +proposal action: promote +proposal live effect: false +historical replay: true +replay reauthorization: false +``` + +The exact point-run evaluation was `impact_evaluation:d94ec6d786c8a4bbfb8038959ca7a7d4` +with material +`sha256:d94ec6d786c8a4bbfb8038959ca7a7d45567d4ddcdb3b6c431b8cdc2dac84c47`. +The exact non-effective proposal was +`impact_governance_proposal:260db0b19aff6fe01cb8f0fbc81f3327` with material +`sha256:260db0b19aff6fe01cb8f0fbc81f33278392794d8652a41442751bfebe79141f`. + +These point-run identities deliberately include exact record-availability and reviewed-action time. +Historical replay in the same durable store returns those exact identities without reauthorization. +Fresh hosts reproduce the source content identities, target definition, probabilities, event +outcome, Brier material, classification, and proposal semantics; they do not pretend independent +wall clocks are the same immutable availability coordinate. + +## Verification + +The frozen World dependency versions plus the stacked Core candidate and separately packaged +reference action adapter produced: + +```text +python -B -m pytest domain_packs/tests/test_p2c3_measured_feedback.py \ + domain_packs/tests/test_p2c4_reviewed_impact_disposition.py \ + domain_packs/tests/test_p2c5_citation_correctness_outcome.py \ + domain_packs/tests/test_p2c6_contradiction_attention_outcome.py \ + domain_packs/tests/test_p2c7_correction_detection_delay_outcome.py \ + domain_packs/tests/test_p2c8_correction_revision_stability_outcome.py \ + domain_packs/tests/test_p2c9_forecast_calibration_outcome.py -q --tb=short +25 passed in 8.59s + +python -B -m pytest -q --tb=short +108 passed in 22.22s + +python -B -m pytest adapters/federal_register_source/tests -q --tb=short +26 passed in 0.22s + +python -B -m pytest tests/test_release_contract.py -q --tb=short +7 passed in 0.04s + +# Installed public ace-core==0.5.0; candidate tests skip explicitly. +python -B -m pytest -q --tb=short -rs +83 passed, 25 skipped in 13.72s + +ruff check --no-cache +PASS + +ruff format --check --no-cache +4 files already formatted + +uv build --out-dir /tmp/ace-p2c9-dist-20260810 +Successfully built unchanged 0.9.0 source distribution and inert data-only wheel + +git diff --check +PASS +``` + +The twenty-five public-Core skips are explicit candidate boundaries: P2C3 through P2C9 require +unreleased stacked Core measured-impact contracts. The public P2C2 journey and every released +boundary remain green. The wheel contains 45 inert Domain Pack JSON/metadata files, no Python or +entry points, and retains `ace-core>=0.5.0,<0.6`. + +Repository-wide Ruff remains an inherited release-hygiene blocker. With the same locked Ruff +version, both the P2C8 parent and this P2C9 worktree report exactly 14 lint findings and 12 format +targets. Scoped P2C9 checks and `git diff --check` are green. This packet does not rewrite unrelated +history, but release closeout must reconcile the repository-wide gate before publication. + +## Claim boundary + +The World forecast contract forbids extra result material and requires every basis reference to be +available at issuance. The review exact-loads the forecast, original Observation, and later +correction Observation, derives the binary event from their explicit correction link, and derives +the Brier contribution from probability and event outcome. The Core Outcome points to that exact +review. Historical replay requires no new authority, and the proposal remains non-effective, +non-selectable, and unapplied. + +This is a single-event probabilistic score over a held-out recorded result. It is not an empirical +calibration curve, population reliability estimate, historically contemporaneous forecast, ACE +probability-generation proof, model-skill finding, live-monitoring result, source-independence +finding, causal estimate, general Brief-quality score, legal-truth claim, or human-benefit result. + +## Remaining work + +A materially different real source, independent Market reproduction, combined-main review/CI, +public Core artifacts, repository-wide lint/format reconciliation, security/release checks, and +opt-in live transport remain future bounded work. Core issue #49 F1, F3, and F5 still require +explicit 0.6 release-owner disposition; this World packet neither implements nor re-dates them. diff --git a/docs/design/world-intelligence-p2c9-forecast-calibration-outcome-work-packet-v1.md b/docs/design/world-intelligence-p2c9-forecast-calibration-outcome-work-packet-v1.md new file mode 100644 index 0000000..e476464 --- /dev/null +++ b/docs/design/world-intelligence-p2c9-forecast-calibration-outcome-work-packet-v1.md @@ -0,0 +1,124 @@ +# World Intelligence P2C9 forecast-calibration outcome work packet (v1) + +**Status:** stacked source-checkout candidate; this packet does not release World Intelligence, +apply a governance proposal, close ACE Core issue #38, pass SI4, or complete ACE 0.6.0. + +**Frozen:** 2026-08-10 from World P2C8 commit +`129767d27d4af22dba292deb9e691375a0695bb8`, stacked on the Core exact observed-result +provenance candidate `433e3d16c5458c975557dcd1552824fb959d4d12`. + +## Objective + +Prove that a World-owned probabilistic forecast can be issued from exact admitted evidence before +an exact result is available, resolved under an inspectable scoring rule, and compared through the +unchanged domain-neutral measured-impact contract: + +```text +original Observation -> treatment probability / probability control + -> Decision -> reviewed Action + -> later correction Observation -> exact resolution review + -> observed Outcome -> useful / harmful / unproven evaluation -> proposal only +``` + +The real source pair remains FCC Federal Register document `2020-28779` and its explicit correction +`2021-10670`. P2C9 intercepts the recorded replay after the original Observation is admitted but +before the correction Observation is appended. It records two exact World forecast artifacts and +completes their reviewed Actions first. Only then does the existing P2C7/P2C8 journey admit and use +the correction result. + +## Frozen forecast and scoring policy + +Both forecasts use the same exact original Observation, target event definition, resolution rule, +window, reviewed workflow, later correction result, and matched conditions. The only deliberate +difference is probability: treatment declares `0.75`; control declares `0.25`. The binary event is +`1.0` only when the later exact admitted source explicitly names the basis document as corrected. + +World owns `world_recorded_binary_forecast_brier_quality` version `candidate-1`. For each exact +forecast/result pair it derives: + +```text +brier_loss = (forecast_probability - binary_event_outcome) ** 2 +brier_quality = 1 - brier_loss +``` + +The observed event is true, so treatment quality is `0.9375`, control quality is `0.4375`, and the +paired effect is `0.5` across two replicated reviewed workflows. The product criterion requires two +matched pairs and a useful threshold of `0.5`. + +This is one single-event Brier contribution. It exercises exact probability/result provenance and +score sensitivity; it is not an empirical calibration curve, population reliability estimate, or +evidence that ACE generated a skillful forecast. The two probabilities are declared fixture inputs, +not model outputs. + +## Exact acceptance + +P2C9 must: + +1. rerun P2C2 through P2C8 and preserve every prior immutable result and proposal; +2. append treatment and control forecast records from the exact original Observation without any + correction identity or correction material in either forecast; +3. complete two reviewed treatment Actions and two reviewed control Actions before the correction + result becomes available; +4. admit the existing exact correction Observation only after those forecast Actions; +5. exact-load the forecast, basis Observation, and correction Observation and derive the binary + outcome from the correction's explicit link to the original document; +6. append four independently authenticated resolution reviews and four Core Outcomes naming those + exact reviews as their observed results; +7. derive, never accept, the Brier loss and quality score from each exact probability/outcome pair; +8. classify the two-pair difference `useful` and emit only a non-effective, non-selectable + `promote` proposal requiring separate human review; +9. replay without reauthorization and reproduce the target definition, probabilities, event + outcome, scores, classification, and proposal semantics across fresh hosts; and +10. reject future/unavailable basis material, result coordinates injected into forecast material, + missing withholding, out-of-range probabilities, changed event linkage, and invented scores. + +## Negative and leakage controls + +The forecast contract has no observed-result field. Its only exact evidence coordinates are the +basis Observations, all of which must be available by forecast issuance. Extra result material is +forbidden. The resolution review requires the result reference to become available after the exact +forecast reference and within the declared window. The harness additionally proves every reviewed +forecast Action completed before the correction record became available. + +The control is not a source-only or no-action baseline. It is an exact probability control under +the same event and workflow. This isolates sensitivity to declared probability while holding result +identity, source linkage, action topology, review policy, and conditions constant. Stacked Core +tests remain authoritative for missing attribution, condition mismatch, cutoff leakage, unavailable +Outcomes, duplicate/replayed evidence, interruption, restart, and denied authority. + +## Ownership boundary + +World owns the public-event target, forecast vocabulary, probability fixtures, binary resolution +rule, Brier-quality mapping, source-policy limits, tests, and evidence. Core owns immutable records, +provenance, Decisions, reviewed Actions, Outcomes, authority, replay, and append-only history. +Intelligence owns domain-neutral conditions, matched evaluation, uncertainty, classification, and +proposal contracts. No FCC, Federal Register, forecast-policy, or Brier noun moves into Core or +Intelligence. + +## Files and rollback + +This packet owns: + +- `scripts/p2c9_forecast_calibration_outcome.py`; +- `domain_packs/tests/test_p2c9_forecast_calibration_outcome.py`; +- the additive P2C7/P2C8 pre-correction state handoff; +- this work packet, its audit, and restrained README/roadmap references. + +It changes no shipped Domain Pack, connector, fixture source policy, package version, dependency +range, lockfile, release record, Core contract, or public artifact. Rollback removes the P2C9 +harness, tests, handoff, and candidate documentation. Forecast, review, Outcome, evaluation, and +proposal records already persisted by a host remain immutable history. + +## Non-claims and next packet + +This packet does not establish a historically contemporaneous forecast, probability generation by +ACE, model skill, population calibration, live monitoring, network freshness, source independence, +causal benefit, legal truth, general Brief quality, or human usefulness. It does not apply the +proposal or grant authority to a Domain Pack. + +The next bounded packet should repeat a correction/outcome measure over a materially different +source or reproduce the unchanged measured-impact contracts independently in Market Intelligence. +Public Core artifacts, combined-main review/CI, compatibility, security and release gates, opt-in +live transport, repository-wide hygiene, and any separately authorized proposal application remain +separate work. Core issue #49 F1, F3, and F5 still require explicit 0.6 release-owner disposition; +this packet does not implement or re-date them. diff --git a/domain_packs/tests/test_p2c9_forecast_calibration_outcome.py b/domain_packs/tests/test_p2c9_forecast_calibration_outcome.py new file mode 100644 index 0000000..20339f2 --- /dev/null +++ b/domain_packs/tests/test_p2c9_forecast_calibration_outcome.py @@ -0,0 +1,200 @@ +from __future__ import annotations + +import copy +import importlib.util +import json + +import pytest +from pydantic import ValidationError + + +def _require_candidate_contracts() -> None: + if importlib.util.find_spec("ace.application.measured_impact") is None: + pytest.skip("P2C9 requires the stacked ACE Core measured-impact candidate") + from ace.intelligence import ImpactOutcomeMeasuresV1Alpha1 + + if "observed_result" not in ImpactOutcomeMeasuresV1Alpha1.model_fields: + pytest.skip("P2C9 requires exact observed-result provenance from the stacked Core candidate") + if importlib.util.find_spec("ace_reference_workspace_action") is None: + pytest.skip("P2C9 requires the separately packaged Core reference adapter") + + +@pytest.mark.asyncio +async def test_exact_withheld_forecast_result_becomes_a_measured_outcome(tmp_path) -> None: + _require_candidate_contracts() + from scripts.p2c9_forecast_calibration_outcome import run_forecast_calibration_outcome + + result = await run_forecast_calibration_outcome(tmp_path) + + assert result["review_policy"]["reviewer_ref"] == "principal:world-forecast-calibration-reviewer" + assert result["evaluation"]["criterion"]["requires_observed_result"] is True + assert result["evaluation"]["classification"] == "useful" + assert result["evaluation"]["matched_pair_count"] == 2 + assert result["evaluation"]["mean_effect"] == 0.5 + assert result["proposal"]["action"] == "promote" + assert result["proposal"]["live_effect"] is False + assert result["proposal"]["selectable"] is False + assert result["proposal"]["requires_human_review"] is True + assert result["replay"] == { + "historical": True, + "no_reauthorization": True, + "transaction_receipt_id": result["replay"]["transaction_receipt_id"], + } + + +@pytest.mark.asyncio +async def test_forecast_material_excludes_the_later_exact_result_and_conditions_match(tmp_path) -> None: + _require_candidate_contracts() + from scripts.p2c9_forecast_calibration_outcome import run_forecast_calibration_outcome + + result = await run_forecast_calibration_outcome(tmp_path) + treatment = result["forecasts"]["treatment"] + control = result["forecasts"]["control"] + reviews = (*result["observed_results"]["treatment"], *result["observed_results"]["control"]) + + assert treatment["basis_observations"] == control["basis_observations"] + assert treatment["target_event_key"] == control["target_event_key"] + assert treatment["target_event_definition_json"] == control["target_event_definition_json"] + assert treatment["policy_digest"] == control["policy_digest"] + assert treatment["probability"] == 0.75 + assert control["probability"] == 0.25 + assert "2021-10670" not in json.dumps(treatment, sort_keys=True) + assert "2021-10670" not in json.dumps(control, sort_keys=True) + assert all(item["result_withheld_until_after_forecast"] for item in reviews) + assert all(item["reviewed_forecast"]["available_at"] < item["observed_result"]["available_at"] for item in reviews) + assert result["source_event"]["result_available_after_forecasts"] is True + assert result["source_event"]["result_available_after_reviewed_actions"] is True + assert ( + result["source_event"]["latest_forecast_action_completed_at"] + < result["source_event"]["observed_result"]["available_at"] + ) + + +@pytest.mark.asyncio +async def test_single_event_brier_contribution_is_derived_not_asserted(tmp_path) -> None: + _require_candidate_contracts() + from scripts.p2c9_forecast_calibration_outcome import run_forecast_calibration_outcome + + result = await run_forecast_calibration_outcome(tmp_path) + treatment = result["observed_results"]["treatment"] + control = result["observed_results"]["control"] + + assert {item["event_outcome"] for item in (*treatment, *control)} == {1.0} + assert {item["explicit_correction_linkage_verified"] for item in (*treatment, *control)} == {True} + assert {item["brier_loss"] for item in treatment} == {0.0625} + assert {item["brier_quality_score"] for item in treatment} == {0.9375} + assert {item["brier_loss"] for item in control} == {0.5625} + assert {item["brier_quality_score"] for item in control} == {0.4375} + assert result["scope"]["population_calibration_claimed"] is False + assert result["scope"]["probability_generated_by_ace_claimed"] is False + assert result["scope"]["historically_contemporaneous_forecast_claimed"] is False + + +@pytest.mark.asyncio +async def test_leaked_basis_missing_withholding_and_invented_brier_material_fail_closed(tmp_path) -> None: + _require_candidate_contracts() + from scripts.p2c9_forecast_calibration_outcome import ( + ForecastResolutionReviewV1Alpha1, + PublicEventForecastV1Alpha1, + run_forecast_calibration_outcome, + ) + + result = await run_forecast_calibration_outcome(tmp_path) + forecast = copy.deepcopy(result["forecasts"]["treatment"]) + forecast["probability"] = 1.1 + forecast["forecast_id"] = None + forecast["forecast_digest"] = None + with pytest.raises(ValidationError): + PublicEventForecastV1Alpha1.model_validate_json(json.dumps(forecast)) + + forecast = copy.deepcopy(result["forecasts"]["treatment"]) + forecast["basis_observations"][0]["available_at"] = forecast["resolution_window_end"] + forecast["forecast_id"] = None + forecast["forecast_digest"] = None + with pytest.raises(ValidationError, match="basis includes evidence unavailable"): + PublicEventForecastV1Alpha1.model_validate_json(json.dumps(forecast)) + + forecast = copy.deepcopy(result["forecasts"]["treatment"]) + forecast["observed_result"] = result["source_event"]["observed_result"] + forecast["forecast_id"] = None + forecast["forecast_digest"] = None + with pytest.raises(ValidationError): + PublicEventForecastV1Alpha1.model_validate_json(json.dumps(forecast)) + + review = copy.deepcopy(result["observed_results"]["treatment"][0]) + review["brier_quality_score"] = 1.0 + review["review_id"] = None + review["review_digest"] = None + with pytest.raises(ValidationError, match="Brier material differs"): + ForecastResolutionReviewV1Alpha1.model_validate_json(json.dumps(review)) + + review = copy.deepcopy(result["observed_results"]["treatment"][0]) + review["result_withheld_until_after_forecast"] = False + review["review_id"] = None + review["review_digest"] = None + with pytest.raises(ValidationError, match="result was not withheld"): + ForecastResolutionReviewV1Alpha1.model_validate_json(json.dumps(review)) + + +@pytest.mark.asyncio +async def test_forecast_material_scores_and_classification_are_deterministic_across_fresh_hosts( + tmp_path, +) -> None: + _require_candidate_contracts() + from scripts.p2c9_forecast_calibration_outcome import run_forecast_calibration_outcome + + first_root = tmp_path / "first" + second_root = tmp_path / "second" + first_root.mkdir() + second_root.mkdir() + first = await run_forecast_calibration_outcome(first_root) + second = await run_forecast_calibration_outcome(second_root) + + assert first["source_event"]["fixture_id"] == second["source_event"]["fixture_id"] + assert first["source_event"]["fixture_digest"] == second["source_event"]["fixture_digest"] + for coordinate in ("original_observation", "observed_result"): + for field in ("record_key", "payload_contract"): + assert first["source_event"][coordinate][field] == second["source_event"][coordinate][field] + for variant in ("treatment", "control"): + for field in ( + "target_event_key", + "target_event_definition_json", + "probability", + "policy_id", + "policy_version", + "limitations", + ): + assert first["forecasts"][variant][field] == second["forecasts"][variant][field] + first_scores = [ + ( + item["forecast_probability"], + item["event_outcome"], + item["brier_loss"], + item["brier_quality_score"], + item["explicit_correction_linkage_verified"], + item["result_withheld_until_after_forecast"], + ) + for item in first["observed_results"][variant] + ] + second_scores = [ + ( + item["forecast_probability"], + item["event_outcome"], + item["brier_loss"], + item["brier_quality_score"], + item["explicit_correction_linkage_verified"], + item["result_withheld_until_after_forecast"], + ) + for item in second["observed_results"][variant] + ] + assert first_scores == second_scores + for field in ( + "classification", + "matched_pair_count", + "mean_effect", + "treatment_mean", + "control_mean", + ): + assert first["evaluation"][field] == second["evaluation"][field] + for field in ("action", "live_effect", "selectable", "requires_human_review"): + assert first["proposal"][field] == second["proposal"][field] diff --git a/scripts/p2c7_correction_detection_delay_outcome.py b/scripts/p2c7_correction_detection_delay_outcome.py index 3c99b27..8c5bf0d 100644 --- a/scripts/p2c7_correction_detection_delay_outcome.py +++ b/scripts/p2c7_correction_detection_delay_outcome.py @@ -4,6 +4,7 @@ import asyncio import json +from collections.abc import Awaitable, Callable from datetime import datetime from pathlib import Path from typing import Any, Literal, Self @@ -548,6 +549,11 @@ async def run_correction_detection_delay_outcome( workspace_root: Path, *, state_sink: dict[str, Any] | None = None, + before_correction: Callable[ + [dict[str, Any], dict[str, Any], ObservationV1Alpha1, ImmutableRecordReferenceV1], + Awaitable[None], + ] + | None = None, ) -> dict[str, Any]: """Run P2C7 over an exact recorded official correction pair.""" @@ -556,6 +562,8 @@ async def run_correction_detection_delay_outcome( environment = state["environment"] fixture = load_correction_fixture() original, original_ref = await _append_source_observation(state, fixture=fixture, role="original") + if before_correction is not None: + await before_correction(state, fixture, original, original_ref) correction, correction_ref = await _append_source_observation(state, fixture=fixture, role="correction") replay = fixture["recorded_replay"] treatment_artifact, treatment_ref, treatment_content = await _append_correction_artifact( diff --git a/scripts/p2c8_correction_revision_stability_outcome.py b/scripts/p2c8_correction_revision_stability_outcome.py index 90a2e32..5aec053 100644 --- a/scripts/p2c8_correction_revision_stability_outcome.py +++ b/scripts/p2c8_correction_revision_stability_outcome.py @@ -573,11 +573,20 @@ def _install_policy(state: dict[str, Any]): return criterion_head, binding -async def run_correction_revision_stability_outcome(workspace_root: Path) -> dict[str, Any]: +async def run_correction_revision_stability_outcome( + workspace_root: Path, + *, + state_sink: dict[str, Any] | None = None, + before_correction=None, +) -> dict[str, Any]: """Run P2C8 over exact prior, stable revision, and drift-control Briefs.""" - state: dict[str, Any] = {} - prior_packet = await run_correction_detection_delay_outcome(workspace_root, state_sink=state) + state: dict[str, Any] = {} if state_sink is None else state_sink + prior_packet = await run_correction_detection_delay_outcome( + workspace_root, + state_sink=state, + before_correction=before_correction, + ) environment = state["environment"] material = _build_briefs(state) prior_ref = await _append_brief(state, brief=material["prior"], role="prior") diff --git a/scripts/p2c9_forecast_calibration_outcome.py b/scripts/p2c9_forecast_calibration_outcome.py new file mode 100644 index 0000000..047afab --- /dev/null +++ b/scripts/p2c9_forecast_calibration_outcome.py @@ -0,0 +1,790 @@ +"""Measure one withheld-result probabilistic forecast under a World-owned rule.""" + +from __future__ import annotations + +import asyncio +import json +from datetime import datetime +from pathlib import Path +from typing import Any, Literal, Self + +from ace.application import MeasuredImpactService +from ace.core import ( + AuthenticatedRuntimeContextV1Alpha1, + CapabilityArtifactIdentityV1Alpha1, + GovernedOperationBindingV1Alpha1, + GovernedStateHeadPreconditionV1Alpha1, + ImmutableRecordReferenceV1, + OutcomeIntentV1Alpha1, + OutcomeV1Alpha1, + canonical_json, + capability_state_ref_for_artifact, +) +from ace.intelligence import ( + ImpactClassification, + ImpactConditionsV1Alpha1, + ImpactCriterionV1Alpha1, + ImpactEvaluationRequestV1Alpha1, + ImpactEvidenceV1Alpha1, + ImpactGovernanceAction, + ImpactMetricDirection, + ImpactOutcomeMeasuresV1Alpha1, + ImpactTargetKind, + ObservationV1Alpha1, +) +from pydantic import Field, field_validator, model_validator + +from scripts.p2c2_federal_register_monitor import _time +from scripts.p2c2_governed_reality_brief import _context, _head +from scripts.p2c3_measured_feedback import ( + ReviewedExport, + _append_value, + _authorize_append, + _digest, + _record_attribution, + _run_reviewed_export, +) +from scripts.p2c5_citation_correctness_outcome import _derive_identity, _FrozenModel +from scripts.p2c7_correction_detection_delay_outcome import ( + _load_observation, + correction_fixture_digest, +) +from scripts.p2c8_correction_revision_stability_outcome import ( + run_correction_revision_stability_outcome, +) + +OUTCOME_TYPE = "independent_artifact_review" +MEASURE_ID = "recorded_binary_forecast_brier_quality" +CRITERION_ID = "impact_criterion:world-recorded-binary-forecast-brier-quality" +REVIEW_POLICY_ID = "world_recorded_binary_forecast_brier_quality" +REVIEW_POLICY_VERSION = "candidate-1" + +RESOLUTION_WINDOW_END = _time("2026-08-11T19:00:00Z") + +TARGET_EVENT_KEY = "recorded-replay:explicit-correction:2020-28779" +TARGET_EVENT_DEFINITION = canonical_json( + { + "basis_document_number": "2020-28779", + "event_type": "later_admitted_source_explicitly_corrects_basis_document", + "resolution_rule": ("true only when a later exact admitted source names the basis document as corrected"), + "source_family": "federal_register", + "withheld_result_policy": "forecast material contains no outcome identity or outcome material", + } +) + +IMPACT_ARTIFACT = CapabilityArtifactIdentityV1Alpha1( + capability="measured_impact_evaluation", + contract="ace.application.measured-impact-service/v1alpha1", + implementation_id="world_recorded_binary_forecast_brier_candidate", + implementation_version="0.1.0", + artifact_digest="sha256:" + "c" * 64, +) + + +class PublicEventForecastV1Alpha1(_FrozenModel): + """One World-owned probability issued before an exact result is available.""" + + contract: Literal["ace.world-intelligence.public-event-forecast/v1alpha1"] = ( + "ace.world-intelligence.public-event-forecast/v1alpha1" + ) + product_id: str + forecast_key: str + basis_observations: tuple[ImmutableRecordReferenceV1, ...] = Field(min_length=1, max_length=16) + target_event_key: str + target_event_definition_json: str + probability: float = Field(ge=0.0, le=1.0) + policy_id: str + policy_version: str + policy_digest: str + issued_at: datetime + resolution_window_start: datetime + resolution_window_end: datetime + limitations: tuple[str, ...] + forecast_id: str | None = None + forecast_digest: str | None = None + + @field_validator("basis_observations") + @classmethod + def canonicalize_basis( + cls, value: tuple[ImmutableRecordReferenceV1, ...] + ) -> tuple[ImmutableRecordReferenceV1, ...]: + ordered = tuple(sorted(value, key=lambda item: item.storage_id)) + if len({item.storage_id for item in ordered}) != len(ordered): + raise ValueError("forecast basis cannot amplify duplicate exact records") + return ordered + + @field_validator("target_event_definition_json") + @classmethod + def validate_target_definition(cls, value: str) -> str: + try: + material = json.loads(value) + except (json.JSONDecodeError, TypeError, ValueError) as exc: + raise ValueError("forecast target definition must be canonical JSON") from exc + if canonical_json(material) != value or material != json.loads(TARGET_EVENT_DEFINITION): + raise ValueError("forecast target definition differs from the frozen World event rule") + return value + + @field_validator("limitations") + @classmethod + def canonicalize_limitations(cls, value: tuple[str, ...]) -> tuple[str, ...]: + ordered = tuple(sorted(value)) + if not ordered or len(ordered) != len(set(ordered)): + raise ValueError("forecast limitations must be non-empty and unique") + return ordered + + @model_validator(mode="after") + def validate_scope_time_and_identity(self) -> Self: + if any(item.product_id != self.product_id for item in self.basis_observations): + raise ValueError("forecast basis crossed exact product scope") + if any(item.available_at > self.issued_at for item in self.basis_observations): + raise ValueError("forecast basis includes evidence unavailable when the forecast was issued") + if not self.issued_at <= self.resolution_window_start < self.resolution_window_end: + raise ValueError("forecast resolution window must begin at or after issuance") + if self.target_event_key != TARGET_EVENT_KEY: + raise ValueError("forecast target key differs from the frozen World event") + _derive_identity( + self, + prefix="public_event_forecast", + id_field="forecast_id", + digest_field="forecast_digest", + ) + return self + + +class ForecastResolutionReviewV1Alpha1(_FrozenModel): + """Exact product-owned resolution and single-event Brier contribution.""" + + contract: Literal["ace.world-intelligence.forecast-resolution-review/v1alpha1"] = ( + "ace.world-intelligence.forecast-resolution-review/v1alpha1" + ) + product_id: str + review_key: str + reviewed_forecast: ImmutableRecordReferenceV1 + basis_observation: ImmutableRecordReferenceV1 + observed_result: ImmutableRecordReferenceV1 + reviewer_context: AuthenticatedRuntimeContextV1Alpha1 + policy_id: str + policy_version: str + policy_digest: str + source_fixture_digest: str + target_event_key: str + target_event_definition_json: str + forecast_issued_at: datetime + resolution_window_start: datetime + resolution_window_end: datetime + forecast_probability: float = Field(ge=0.0, le=1.0) + result_withheld_until_after_forecast: bool + explicit_correction_linkage_verified: bool + event_outcome: float = Field(ge=0.0, le=1.0) + brier_loss: float = Field(ge=0.0, le=1.0) + brier_quality_score: float = Field(ge=0.0, le=1.0) + limitations: tuple[str, ...] + reviewed_at: datetime + review_id: str | None = None + review_digest: str | None = None + + @field_validator("target_event_definition_json") + @classmethod + def validate_target_definition(cls, value: str) -> str: + if value != TARGET_EVENT_DEFINITION: + raise ValueError("forecast review changed the frozen target event definition") + return value + + @field_validator("limitations") + @classmethod + def canonicalize_limitations(cls, value: tuple[str, ...]) -> tuple[str, ...]: + ordered = tuple(sorted(value)) + if not ordered or len(ordered) != len(set(ordered)): + raise ValueError("forecast review limitations must be non-empty and unique") + return ordered + + @model_validator(mode="after") + def validate_scope_resolution_score_and_identity(self) -> Self: + references = (self.reviewed_forecast, self.basis_observation, self.observed_result) + if any(item.product_id != self.product_id for item in references): + raise ValueError("forecast resolution review crossed exact product scope") + if self.reviewed_forecast.storage_id in { + self.basis_observation.storage_id, + self.observed_result.storage_id, + }: + raise ValueError("forecast, basis, and observed result require distinct exact records") + withheld = self.observed_result.available_at > self.reviewed_forecast.available_at + if self.result_withheld_until_after_forecast != withheld or not withheld: + raise ValueError("forecast result was not withheld until after exact forecast availability") + if not ( + self.forecast_issued_at + <= self.resolution_window_start + <= self.observed_result.available_at + <= self.reviewed_at + <= self.resolution_window_end + ): + raise ValueError("forecast resolution escaped its declared time window") + if self.target_event_key != TARGET_EVENT_KEY: + raise ValueError("forecast review changed the frozen target event") + expected_outcome = float(self.explicit_correction_linkage_verified) + if self.event_outcome != expected_outcome: + raise ValueError("forecast outcome differs from exact correction linkage") + expected_loss = (self.forecast_probability - self.event_outcome) ** 2 + expected_score = 1.0 - expected_loss + if self.brier_loss != expected_loss or self.brier_quality_score != expected_score: + raise ValueError("forecast Brier material differs from the frozen product formula") + _derive_identity( + self, + prefix="forecast_resolution_review", + id_field="review_id", + digest_field="review_digest", + ) + return self + + +class _ReplayMustNotAuthorize: + async def authorize_action(self, request): + raise AssertionError(f"historical forecast evaluation requested new authority: {request.authorization_key}") + + +def _policy_digest(original_ref: ImmutableRecordReferenceV1, *, issued_at: datetime) -> str: + return _digest( + { + "policy_id": REVIEW_POLICY_ID, + "policy_version": REVIEW_POLICY_VERSION, + "basis_observation": original_ref.model_dump(mode="json"), + "target_event_key": TARGET_EVENT_KEY, + "target_event_definition_json": TARGET_EVENT_DEFINITION, + "resolution_window_start": issued_at.isoformat(), + "resolution_window_end": RESOLUTION_WINDOW_END.isoformat(), + "score": "1 - (forecast_probability - binary_event_outcome) ** 2", + "claim_boundary": "single-event Brier contribution, not population calibration", + } + ) + + +def _install_policy(state: dict[str, Any]): + environment = state["environment"] + runtime = state["runtime"] + product_id = environment.fixture["product_id"] + criterion_head = _head(product_id, "impact_criterion", CRITERION_ID, 100) + operation_head = _head( + product_id, + "governed_operation_configuration", + "governed_operation_configuration:world-recorded-binary-forecast-brier-quality", + 101, + ) + binding = GovernedOperationBindingV1Alpha1( + product_id=product_id, + artifact=IMPACT_ARTIFACT, + configuration_ref=operation_head.state_id, + authority="append_measured_impact", + grant_ref="authority_grant:world-recorded-binary-forecast-brier-quality", + state_head_precondition=GovernedStateHeadPreconditionV1Alpha1.from_head(operation_head), + ) + capability_head = _head( + product_id, + "capability_state", + capability_state_ref_for_artifact(IMPACT_ARTIFACT), + 102, + ) + authority_head = _head(product_id, "authority_grant", binding.grant_ref, 103) + for head in (criterion_head, operation_head, capability_head, authority_head): + environment.store.set_governed_state_head(head) + runtime.heads[head.state_kind, head.state_id] = head + runtime.bindings = (*runtime.bindings, binding) + return criterion_head, binding + + +async def _append_forecast( + state: dict[str, Any], + *, + original_ref: ImmutableRecordReferenceV1, + variant: Literal["treatment", "control"], + probability: float, + issued_at: datetime, +) -> tuple[PublicEventForecastV1Alpha1, ImmutableRecordReferenceV1, str]: + environment = state["environment"] + forecast = PublicEventForecastV1Alpha1( + product_id=environment.fixture["product_id"], + forecast_key=f"recorded-binary-forecast:{variant}:2020-28779", + basis_observations=(original_ref,), + target_event_key=TARGET_EVENT_KEY, + target_event_definition_json=TARGET_EVENT_DEFINITION, + probability=probability, + policy_id=REVIEW_POLICY_ID, + policy_version=REVIEW_POLICY_VERSION, + policy_digest=_policy_digest(original_ref, issued_at=issued_at), + issued_at=issued_at, + resolution_window_start=issued_at, + resolution_window_end=RESOLUTION_WINDOW_END, + limitations=( + "declared_fixture_probability_not_generated_by_ace", + "offline_recorded_replay_not_live_forecasting", + "single_event_brier_contribution_not_population_calibration", + ), + ) + requested_at = state["clock"]() + authorization = await _authorize_append( + state, + context=environment.context, + authorization_key=f"append:world-recorded-binary-forecast:{variant}", + subject_ref=str(forecast.forecast_id), + subject_digest=str(forecast.forecast_digest), + requested_at=requested_at, + ) + reference = await _append_value( + state, + value=forecast, + record_kind="public_event_forecast", + record_key=str(forecast.forecast_id), + transaction_key=f"public-event-forecast:{forecast.forecast_id}", + as_of=forecast.issued_at, + authorization=authorization, + ) + content = ( + "# Recorded Binary Public-Event Forecast\n\n" + f"Basis Observation: {original_ref.storage_id}\n\n" + f"Target: {forecast.target_event_key}\n\n" + f"Declared probability: {forecast.probability}\n\n" + "The exact outcome identity and outcome material are withheld from this forecast record.\n" + ) + return forecast, reference, content + + +async def _issue_forecasts_before_correction( + state: dict[str, Any], + _fixture: dict[str, Any], + original: ObservationV1Alpha1, + original_ref: ImmutableRecordReferenceV1, +) -> None: + original_payload = original.payload.parsed_value()["record"] + if original_payload["document_number"] != "2020-28779": + raise AssertionError("forecast basis changed its exact original public record") + criterion_head, impact_binding = _install_policy(state) + issued_at = state["clock"]() + treatment, treatment_ref, treatment_content = await _append_forecast( + state, + original_ref=original_ref, + variant="treatment", + probability=0.75, + issued_at=issued_at, + ) + control, control_ref, control_content = await _append_forecast( + state, + original_ref=original_ref, + variant="control", + probability=0.25, + issued_at=issued_at, + ) + treatment_exports = tuple( + [ + await _run_reviewed_export( + state, + subject=treatment_ref, + content=treatment_content, + pair_index=index, + variant="forecast-calibration-treatment", + ) + for index in (1, 2) + ] + ) + control_exports = tuple( + [ + await _run_reviewed_export( + state, + subject=control_ref, + content=control_content, + pair_index=index, + variant="forecast-calibration-control", + ) + for index in (1, 2) + ] + ) + state.update( + { + "p2c9_criterion_head": criterion_head, + "p2c9_criterion_frozen_at": issued_at, + "p2c9_forecast_issued_at": issued_at, + "p2c9_impact_binding": impact_binding, + "p2c9_treatment_forecast": treatment, + "p2c9_treatment_forecast_ref": treatment_ref, + "p2c9_control_forecast": control, + "p2c9_control_forecast_ref": control_ref, + "p2c9_treatment_exports": treatment_exports, + "p2c9_control_exports": control_exports, + } + ) + + +async def _load_forecast(state: dict[str, Any], reference: ImmutableRecordReferenceV1) -> PublicEventForecastV1Alpha1: + record = await state["environment"].store.load_record( + reference.storage_id, + product_id=reference.product_id, + record_space=reference.record_space, + record_kind=reference.record_kind, + ) + if ( + record is None + or record.reference() != reference + or record.payload_contract != "ace.world-intelligence.public-event-forecast/v1alpha1" + ): + raise AssertionError("recorded forecast is unavailable or changed") + return PublicEventForecastV1Alpha1.model_validate(record.payload) + + +async def _review_forecast( + state: dict[str, Any], + *, + forecast_ref: ImmutableRecordReferenceV1, + original_ref: ImmutableRecordReferenceV1, + correction_ref: ImmutableRecordReferenceV1, + pair_index: int, + variant: Literal["treatment", "control"], +) -> tuple[ForecastResolutionReviewV1Alpha1, ImmutableRecordReferenceV1]: + environment = state["environment"] + forecast = await _load_forecast(state, forecast_ref) + original = await _load_observation(state, original_ref) + correction = await _load_observation(state, correction_ref) + original_payload = original.payload.parsed_value()["record"] + correction_payload = correction.payload.parsed_value()["record"] + linkage_verified = bool( + correction_payload["corrects_document_number"] == original_payload["document_number"] + and correction_payload["document_number"] == "2021-10670" + and forecast.basis_observations == (original_ref,) + and forecast.target_event_key == TARGET_EVENT_KEY + ) + event_outcome = float(linkage_verified) + brier_loss = (forecast.probability - event_outcome) ** 2 + reviewed_at = state["clock"]() + reviewer = _context(environment.context, "principal:world-forecast-calibration-reviewer") + review = ForecastResolutionReviewV1Alpha1( + product_id=environment.fixture["product_id"], + review_key=f"recorded-binary-forecast-resolution:{pair_index}:{variant}", + reviewed_forecast=forecast_ref, + basis_observation=original_ref, + observed_result=correction_ref, + reviewer_context=reviewer, + policy_id=forecast.policy_id, + policy_version=forecast.policy_version, + policy_digest=forecast.policy_digest, + source_fixture_digest=correction_fixture_digest(state["p2c7_fixture"]), + target_event_key=forecast.target_event_key, + target_event_definition_json=forecast.target_event_definition_json, + forecast_issued_at=forecast.issued_at, + resolution_window_start=forecast.resolution_window_start, + resolution_window_end=forecast.resolution_window_end, + forecast_probability=forecast.probability, + result_withheld_until_after_forecast=correction_ref.available_at > forecast_ref.available_at, + explicit_correction_linkage_verified=linkage_verified, + event_outcome=event_outcome, + brier_loss=brier_loss, + brier_quality_score=1.0 - brier_loss, + limitations=forecast.limitations, + reviewed_at=reviewed_at, + ) + requested_at = state["clock"]() + authorization = await _authorize_append( + state, + context=reviewer, + authorization_key=f"recorded-binary-forecast-resolution:{pair_index}:{variant}", + subject_ref=str(review.review_id), + subject_digest=str(review.review_digest), + requested_at=requested_at, + ) + reference = await _append_value( + state, + value=review, + record_kind="forecast_resolution_review", + record_key=str(review.review_id), + transaction_key=f"forecast-resolution-review:{review.review_id}", + as_of=review.reviewed_at, + authorization=authorization, + ) + return review, reference + + +async def _record_review_outcome( + state: dict[str, Any], + *, + export: ReviewedExport, + review: ForecastResolutionReviewV1Alpha1, + review_ref: ImmutableRecordReferenceV1, + pair_index: int, + variant: Literal["treatment", "control"], +) -> ImmutableRecordReferenceV1: + environment = state["environment"] + latency_ms = int((review.reviewed_at - review.forecast_issued_at).total_seconds() * 1_000) + measures = ImpactOutcomeMeasuresV1Alpha1( + primary_value=review.brier_quality_score, + observed_result=review_ref, + latency_ms=latency_ms, + cost_usd=0.0, + failure_count=0, + degraded=False, + limitations=review.limitations, + ) + recorded_at = state["clock"]() + observer = _context(environment.context, "principal:world-forecast-calibration-outcome-observer") + intent = OutcomeIntentV1Alpha1( + product_id=environment.fixture["product_id"], + authenticated_context=observer, + decision=export.decision_ref, + outcome_type=OUTCOME_TYPE, + measure_id=MEASURE_ID, + value_json=canonical_json(measures.model_dump(mode="json")), + observed_at=review.reviewed_at, + recorded_at=recorded_at, + ) + authorization = await _authorize_append( + state, + context=observer, + authorization_key=f"recorded-binary-forecast-outcome:{pair_index}:{variant}", + subject_ref=str(intent.intent_id), + subject_digest=str(intent.intent_digest), + requested_at=recorded_at, + ) + outcome = OutcomeV1Alpha1(intent=intent, authorization=authorization) + return await _append_value( + state, + value=outcome, + record_kind="outcome", + record_key=str(outcome.outcome_id), + transaction_key=f"outcome:{outcome.outcome_id}", + as_of=review.reviewed_at, + authorization=authorization, + ) + + +async def run_forecast_calibration_outcome(workspace_root: Path) -> dict[str, Any]: + """Run P2C9 over a withheld exact correction result and declared probabilities.""" + + state: dict[str, Any] = {} + prior_packet = await run_correction_revision_stability_outcome( + workspace_root, + state_sink=state, + before_correction=_issue_forecasts_before_correction, + ) + environment = state["environment"] + original_ref = state["p2c7_original_observation_ref"] + correction_ref = state["p2c7_correction_observation_ref"] + treatment_ref = state["p2c9_treatment_forecast_ref"] + control_ref = state["p2c9_control_forecast_ref"] + treatment_exports = state["p2c9_treatment_exports"] + control_exports = state["p2c9_control_exports"] + if correction_ref.available_at <= max(treatment_ref.available_at, control_ref.available_at): + raise AssertionError("held-out correction became available before the exact forecasts") + latest_forecast_action_completed_at = max( + *(item.terminal.result.completed_at for item in treatment_exports), + *(item.terminal.result.completed_at for item in control_exports), + ) + if correction_ref.available_at <= latest_forecast_action_completed_at: + raise AssertionError("held-out correction became available before the reviewed forecast actions completed") + + evidence: list[ImpactEvidenceV1Alpha1] = [] + treatment_reviews: list[ForecastResolutionReviewV1Alpha1] = [] + control_reviews: list[ForecastResolutionReviewV1Alpha1] = [] + for index, (treatment_export, control_export) in enumerate( + zip(treatment_exports, control_exports, strict=True), start=1 + ): + treatment_attribution = await _record_attribution( + state, + export=treatment_export, + subject=treatment_ref, + pair_index=index, + variant="forecast-calibration-treatment", + ) + control_attribution = await _record_attribution( + state, + export=control_export, + subject=control_ref, + pair_index=index, + variant="forecast-calibration-control", + ) + treatment_review, treatment_review_ref = await _review_forecast( + state, + forecast_ref=treatment_ref, + original_ref=original_ref, + correction_ref=correction_ref, + pair_index=index, + variant="treatment", + ) + control_review, control_review_ref = await _review_forecast( + state, + forecast_ref=control_ref, + original_ref=original_ref, + correction_ref=correction_ref, + pair_index=index, + variant="control", + ) + treatment_outcome = await _record_review_outcome( + state, + export=treatment_export, + review=treatment_review, + review_ref=treatment_review_ref, + pair_index=index, + variant="treatment", + ) + control_outcome = await _record_review_outcome( + state, + export=control_export, + review=control_review, + review_ref=control_review_ref, + pair_index=index, + variant="control", + ) + treatment_reviews.append(treatment_review) + control_reviews.append(control_review) + conditions = ImpactConditionsV1Alpha1( + product_id=environment.fixture["product_id"], + condition_key=f"world-recorded-binary-forecast-brier-pair:{index}", + route_id="world:fcc-recorded-binary-forecast-review", + context_json=canonical_json( + { + "basis_observation": original_ref.model_dump(mode="json"), + "pair_index": index, + "review_policy_digest": treatment_review.policy_digest, + "target_event_definition_json": TARGET_EVENT_DEFINITION, + "target_event_key": TARGET_EVENT_KEY, + "task": "score_declared_probability_after_exact_withheld_result", + "withheld_result_policy": "no outcome coordinate in frozen conditions", + } + ), + observation_window_start=state["p2c9_forecast_issued_at"], + observation_window_end=RESOLUTION_WINDOW_END, + frozen_at=state["p2c9_criterion_frozen_at"], + ) + evidence.append( + ImpactEvidenceV1Alpha1( + product_id=environment.fixture["product_id"], + evidence_key=f"world-recorded-binary-forecast-brier:{index}", + treatment_attribution=treatment_attribution, + control_attribution=control_attribution, + treatment_decision=treatment_export.decision_ref, + control_decision=control_export.decision_ref, + treatment_action_review=treatment_export.review_ref, + treatment_action_admission=treatment_export.admission_ref, + treatment_action_terminal=treatment_export.terminal_ref, + control_action_review=control_export.review_ref, + control_action_admission=control_export.admission_ref, + control_action_terminal=control_export.terminal_ref, + treatment_outcome=treatment_outcome, + control_outcome=control_outcome, + treatment_conditions=conditions, + control_conditions=conditions, + ) + ) + + criterion = ImpactCriterionV1Alpha1( + product_id=environment.fixture["product_id"], + criterion_id=CRITERION_ID, + criterion_version="candidate-1", + target_kind=ImpactTargetKind.INTELLIGENCE_ARTIFACT, + outcome_type=OUTCOME_TYPE, + measure_id=MEASURE_ID, + metric_direction=ImpactMetricDirection.HIGHER_IS_BETTER, + useful_effect_threshold=0.5, + harmful_effect_threshold=0.5, + minimum_matched_pairs=2, + requires_observed_result=True, + harmful_action=ImpactGovernanceAction.ROLLBACK, + state_head_precondition=GovernedStateHeadPreconditionV1Alpha1.from_head(state["p2c9_criterion_head"]), + frozen_at=state["p2c9_criterion_frozen_at"], + ) + request = ImpactEvaluationRequestV1Alpha1( + evaluation_key="world-recorded-binary-forecast-brier:fcc-2021-10670", + product_id=environment.fixture["product_id"], + authenticated_context=environment.context, + criterion=criterion, + target=treatment_ref, + control=control_ref, + evidence=tuple(evidence), + cutoff_at=state["clock"](), + requested_at=state["clock"](), + ) + service = MeasuredImpactService( + store=environment.store, + authorizer=state["reasoning"], + operation_binding=state["p2c9_impact_binding"], + ) + admission = await service.evaluate(request) + historical = await MeasuredImpactService( + store=environment.store, + authorizer=_ReplayMustNotAuthorize(), + operation_binding=state["p2c9_impact_binding"], + ).evaluate(request) + if admission.replayed or not historical.replayed or historical.evaluation != admission.evaluation: + raise AssertionError("forecast evaluation did not replay exact historical material") + if admission.evaluation.classification is not ImpactClassification.USEFUL: + raise AssertionError("frozen forecast Brier criterion did not classify useful") + if admission.proposal is None or admission.proposal.action is not ImpactGovernanceAction.PROMOTE: + raise AssertionError("forecast evaluation did not emit its proposal-only mapping") + if {item.brier_quality_score for item in treatment_reviews} != {0.9375}: + raise AssertionError("treatment forecast lost its exact Brier contribution") + if {item.brier_quality_score for item in control_reviews} != {0.4375}: + raise AssertionError("control forecast lost its exact Brier contribution") + + return { + "contract": "ace.world-intelligence.forecast-calibration-outcome/v1alpha1", + "prior_revision_stability": prior_packet, + "source_event": { + "fixture_id": state["p2c7_fixture"]["fixture_id"], + "fixture_digest": correction_fixture_digest(state["p2c7_fixture"]), + "original_observation": original_ref.model_dump(mode="json"), + "observed_result": correction_ref.model_dump(mode="json"), + "result_available_after_forecasts": True, + "latest_forecast_action_completed_at": latest_forecast_action_completed_at.isoformat(), + "result_available_after_reviewed_actions": True, + }, + "review_policy": { + "policy_id": REVIEW_POLICY_ID, + "policy_version": REVIEW_POLICY_VERSION, + "policy_digest": treatment_reviews[0].policy_digest, + "reviewer_ref": treatment_reviews[0].reviewer_context.actor_ref, + "score_rule": "1 - (forecast_probability - binary_event_outcome) ** 2", + "claim_boundary": "single-event Brier contribution, not population calibration", + }, + "forecasts": { + "treatment": state["p2c9_treatment_forecast"].model_dump(mode="json"), + "control": state["p2c9_control_forecast"].model_dump(mode="json"), + }, + "observed_results": { + "treatment": tuple(item.model_dump(mode="json") for item in treatment_reviews), + "control": tuple(item.model_dump(mode="json") for item in control_reviews), + }, + "evaluation": admission.evaluation.model_dump(mode="json"), + "proposal": admission.proposal.model_dump(mode="json"), + "replay": { + "historical": historical.replayed, + "no_reauthorization": True, + "transaction_receipt_id": str(historical.transaction_receipt.receipt_id), + }, + "scope": { + "exact_real_correction_result": True, + "result_withheld_until_after_forecast_records": True, + "exact_probability_and_result_scoring": True, + "network_access": False, + "historically_contemporaneous_forecast_claimed": False, + "probability_generated_by_ace_claimed": False, + "model_forecast_skill_claimed": False, + "population_calibration_claimed": False, + "causality_claimed": False, + "human_benefit_claimed": False, + "proposal_applied": False, + "autonomous_publication": False, + }, + } + + +def main() -> None: + import argparse + + parser = argparse.ArgumentParser() + parser.add_argument("workspace_root", type=Path) + args = parser.parse_args() + print( + json.dumps( + asyncio.run(run_forecast_calibration_outcome(args.workspace_root)), + indent=2, + sort_keys=True, + ) + ) + + +if __name__ == "__main__": + main()