diff --git a/docs/plans/2026-07-30-bilingual-epigraphs-design.md b/docs/plans/2026-07-30-bilingual-epigraphs-design.md new file mode 100644 index 000000000..b4abe84e7 --- /dev/null +++ b/docs/plans/2026-07-30-bilingual-epigraphs-design.md @@ -0,0 +1,30 @@ +# Ranger Archive V2 Bilingual Epigraphs — Design + +## Deliverable + +Create twenty bilingual original epigraphs, one for each V2 project image. Every item contains one self-sufficient English sentence and one self-sufficient Chinese sentence expressing the same idea. + +## Editorial direction + +- These are original epigraphs, not quotations from the referenced science-fiction works. +- English is written as English rather than produced by literal word-for-word translation. +- Chinese may be more compressed, but may not introduce a different claim. +- Each pair must refer to a concrete object or event in its corresponding V2 image. +- Each pair must also express the project’s distinctive act of knowing: measuring, coupling, testing, recovering, certifying, remembering, bounding, transporting, or generating. +- Avoid generic interchangeable claims about humanity being small, the universe being large, seeking answers, or facing the unknown. +- Use no quotation marks or false attribution to an author or fictional character. +- Keep each sentence short enough to sit beneath a PR hero image. + +## Output structure + +The final artifact is `docs/showcase/ranger-archive/CAPTIONS-v2.md`. Items are numbered 01–20 in the same order as the image manifest. Each item includes the issue or project, the science-fiction lens, an English sentence, and a Chinese sentence. + +## Acceptance criteria + +1. Exactly twenty numbered bilingual pairs exist. +2. Every English and Chinese entry is one complete sentence. +3. Every item maps to the same-numbered V2 image and project. +4. No pair is presented as a source quotation. +5. No two pairs are interchangeable without losing a specific visual or scientific connection. +6. English and Chinese agree in meaning while remaining idiomatic. + diff --git a/docs/plans/2026-07-30-lost-scifi-covers-v2-design.md b/docs/plans/2026-07-30-lost-scifi-covers-v2-design.md new file mode 100644 index 000000000..9a6e930ca --- /dev/null +++ b/docs/plans/2026-07-30-lost-scifi-covers-v2-design.md @@ -0,0 +1,65 @@ +# Lost Science-Fiction Covers V2 — Design + +## Purpose + +Redesign the twenty Ranger Archive banners as twenty unrelated “lost science-fiction book covers.” The first version failed because a shared deep-space palette, glowing geometry, circles, networks, and a recurring probe made distinct subjects look like variants of one generic AI-science-fiction image. + +The second version prioritizes human-scale places, used objects, weather, season, and publishing history. It does not retain the Ranger probe or a common rendering style. + +## Approved direction + +- Direction: **twenty books from different eras and publishing cultures**, 100%. +- Human presence: people do not need to appear; their use, absence, labor, and memory must remain visible in objects and environments. +- Variation: each image uses a materially different medium, location, time, weather, palette, and dominant silhouette. +- Recognition: each science-fiction work is recognizable through an iconic but ordinary setting or object. +- Copyright boundary: do not copy a film frame, actor likeness, franchise logo, proprietary spacecraft, or cover layout. Reinterpret the work through original composition. +- Scientific subject: appears as one anomaly inside the ordinary scene, not as a generic glowing diagram. +- Output: 2.5:1 PR banner, no generated title or caption. +- Preservation: keep the first version under `assets/missions`; save the redesign under `assets/missions-v2`. + +## Twenty covers + +| # | Project × work | Publishing medium | Human place and trace | Scientific anomaly | +|---:|---|---|---|---| +| 01 | Issue #148 × *2001: A Space Odyssey* | 1970s large-format landscape photograph | Dewy dawn grassland, survey stakes, abandoned canvas field bag | A severe black slab divides triangular and honeycomb plots; brass measuring rods imply √5 | +| 02 | Issue #35 × *The Wandering Earth* | 1960s Chinese gouache science poster | Snowbound school courtyard, enamel lunch bowl, red wool scarf, engine towers beyond rooftops | Water in the bowl and window glass share the same phonon ripple | +| 03 | Issue #114 × *The Martian* | Faded instant-film field documentation | Mars greenhouse bench, potatoes, duct tape, wrench, stained verification checklist | One repaired tensor device passes a row of physical reference tests | +| 04 | Issue #71 × *The Hitchhiker’s Guide to the Galaxy* | 1980s British comic paperback illustration | Roadside café table, tea, towel, scribbled receipt bearing 42, absurd folded road atlas | All routes run backward from 42 toward an empty question box | +| 05 | Issue #82 × *I, Robot* | 1950s three-color silkscreen | Drafting desk, typewriter, punched rule cards, pencil corrections, articulated metal hand | Three rule cards admit one syntax branch and mechanically reject another | +| 06 | Issue #73 × *Arrival* | Ink wash and translucent watercolor | Foggy meadow outside a temporary language classroom, wet glass, notebook and rubber gloves | A single imperfect ink glyph leaves a measurable phase offset at its seam | +| 07 | Quantum Geometry × *Flatland* | 1880s hand-colored copper engraving | Victorian study, checked tablecloth, ruler, teacup and spectacles | The printed grid physically rises and bends while its ink inhabitants remain flat | +| 08 | Issue #230 × *Anathem* | Medieval illuminated manuscript and egg tempera | Sunlit stone cloister, open proof book, wax tablet, extinguished candle | Shadows of two columns form certified lower and upper bounds around a gold horizon | +| 09 | Issue #115 × *Project Hail Mary* | 1970s technical-airbrush paperback | Improvised interstellar workbench, coffee pouch, burned circuit, handmade alien material sample | A fragile circuit is rebuilt as a rugged rust-colored core across two environmental zones | +| 10 | Issue #33 × *Ender’s Game* | 1980s arcade-box screenprint | Empty dormitory desk, scuffed shoes, handheld tactical board, wooden fleet pieces | One economical gold move solves a board crowded with costly alternatives | +| 11 | Issue #36 × *The Expanse* | Industrial photo-collage and shipping ephemera | Worn asteroid cargo dock, grease pencil, gloves, ration tin, numbered loading bays | A conveyor cycle returns to its start while one sealed parcel advances exactly one bay | +| 12 | Issue #129 × *Interstellar* | Naturalistic 35mm rural cinema still | Green cornfield and grass, weathered farmhouse, porch chalkboard, wristwatch and child’s model | A determinant grid on the board bends toward a small gravitational distortion in the sky | +| 13 | Issue #121 × *Roadside Picnic* | Gritty 1970s Eastern European documentary photograph | Overgrown rail yard, abandoned lunch, survey flags, muddy boot prints | Dangerous debris floats above false paths; one ordinary footprint trail remains sign-safe | +| 14 | Issue #122 × *Solaris* | 1960s Polish art-film still | Quiet dacha room, empty chair, half-cut apple, analog instruments, window onto the sea | The ocean repeats the room’s waveform with a critical delay | +| 15 | Issue #123 × *Recursion* | 1990s family-photo contact sheet and VHS texture | The same kitchen photographed at four dates, tape recorder, birthday objects subtly displaced | An earlier frame leaks into a later one, making memory visibly non-Markovian | +| 16 | Issue #15 × *Death’s End* | Contemporary Chinese literary-cover painting | Golden wheat field on black crystalline soil, old silver jacket hung on a scarecrow | Wheat bends into opposite left- and right-handed gravitational textures | +| 17 | Hydrodynamics × *Dune* | Japanese woodblock print with mineral pigments | Desert survey camp, water flask, folded cloth, footprints erased by wind | Individual grains transition into river-like streamlines and one breaking vortex | +| 18 | Issue #158 × *The Dark Forest* | Rough charcoal and linocut | Snowy forest clearing, cold camp stove, radio notebook, distant dark trunks | Sparse local tree marks gradually align into long-range order across the whole forest | +| 19 | Issue #128 × *Foundation* | 1950s modernist paper collage | Vast reading room under a city dome, paper timelines, punched cards, librarian’s lamp | Many future strips diverge; a transparent error ruler contains only certified trajectories | +| 20 | Issue #133 × *The Last Question* | 1960s risograph educational-magazine cover | Empty classroom at sunrise, chalk dust, stacks of accepted and rejected question cards | An electromechanical problem press validates many inputs and releases one blank luminous seed | + +## Anti-repetition matrix + +- No recurring protagonist, probe, logo, border, or color accent. +- Only #06 may use a circular glyph as the dominant form. +- Only #11 may use a closed transport cycle, rendered as a physical conveyor rather than a light ring. +- Only #12 may feature a gravitational sky phenomenon. +- Only #18 may use a predominantly black forest. +- At least twelve covers are daylight or interior-light scenes rather than deep space or night. +- At least ten covers use visibly worn domestic, agricultural, educational, or workshop objects. +- Adjacent covers must differ in medium, dominant color, horizon geometry, and spatial scale. + +## Acceptance criteria + +1. Twenty distinct V2 PNG files exist, one per mapped project. +2. Every file is approximately 2.5:1 and suitable for a PR banner. +3. No two files are byte-identical or near-duplicate compositions. +4. A contact sheet shows immediate medium, palette, location, and silhouette variation. +5. Each image contains the specified recognizable setting/object and scientific anomaly. +6. No image contains visible franchise logos, copied actors, watermarks, or unusable generated text. +7. The first version remains intact for comparison. + diff --git a/docs/showcase/ranger-archive/CAPTIONS-v2.md b/docs/showcase/ranger-archive/CAPTIONS-v2.md new file mode 100644 index 000000000..b10f1cf08 --- /dev/null +++ b/docs/showcase/ranger-archive/CAPTIONS-v2.md @@ -0,0 +1,124 @@ +# Ranger Archive V2 — Bilingual Original Epigraphs + +以下文字均为原创题记,不是对应科幻作品的原文引语。编号与 V2 图片及项目清单一致。 + +### 01 — Issue #148 × *2001: A Space Odyssey* + +**EN** Before the ratio had a name, morning dew had already settled on the ruler. + +**中文** 在那个比例获得名字之前,晨露早已落在尺上。 + +### 02 — Issue #35 × *The Wandering Earth* + +**EN** A planet begins to move when its smallest vibrations find something ordinary enough to carry them. + +**中文** 当最微小的振动找到一件日常之物承载它,整颗行星便开始移动。 + +### 03 — Issue #114 × *The Martian* + +**EN** On Mars, trust is not a promise but the last green lamp after every failed test. + +**中文** 在火星上,可信不是一句承诺,而是所有失败测试之后最后亮起的绿灯。 + +### 04 — Issue #71 × *The Hitchhiker’s Guide to the Galaxy* + +**EN** The answer was waiting on the table; we only had to unfold the road back to the question. + +**中文** 答案一直等在桌上;我们只需展开那条返回问题的路。 + +### 05 — Issue #82 × *I, Robot* + +**EN** A language becomes trustworthy when it knows how to reject a beautiful mistake. + +**中文** 一种语言懂得拒绝漂亮的错误时,才真正值得信任。 + +### 06 — Issue #73 × *Arrival* + +**EN** We returned to the same point, but the rain on the glass remembered the journey. + +**中文** 我们回到了同一点,玻璃上的雨却记得这段旅程。 + +### 07 — Quantum Geometry × *Flatland* + +**EN** The world seemed flat only because no one had lifted the tablecloth. + +**中文** 世界之所以显得平坦,只因从未有人掀起那块桌布。 + +### 08 — Issue #230 × *Anathem* + +**EN** Proof is the narrow strip of sunlight that survives between two honest shadows. + +**中文** 证明,是两道诚实阴影之间幸存的那一线阳光。 + +### 09 — Issue #115 × *Project Hail Mary* + +**EN** Between two stars, a tool becomes real when it can be rebuilt from what both worlds leave on the bench. + +**中文** 在两颗恒星之间,能用两个世界留在工作台上的材料重新造出的工具,才真正存在。 + +### 10 — Issue #33 × *Ender’s Game* + +**EN** The efficient move is not the one that avoids the board, but the one that finally understands it. + +**中文** 高效的一步并非绕开棋盘,而是终于看懂棋盘。 + +### 11 — Issue #36 × *The Expanse* + +**EN** Everything returned to where it began—except the parcel the cycle was built to carry. + +**中文** 一切都回到了起点——除了这个周期本就要运走的货箱。 + +### 12 — Issue #129 × *Interstellar* + +**EN** Beyond the cornfield, the sky bent; on the porch, someone had already begun to calculate. + +**中文** 玉米地之外,天空开始弯曲;门廊之上,计算早已开始。 + +### 13 — Issue #121 × *Roadside Picnic* + +**EN** The safe path through the Zone was never the brightest, only the one that left ordinary footprints. + +**中文** 穿过禁区的安全之路从来不是最明亮的,只是那条留下普通脚印的路。 + +### 14 — Issue #122 × *Solaris* + +**EN** We watched the ocean for a reply until it learned the waveform of our question. + +**中文** 我们凝视海洋等待回应,直到它学会了问题的波形。 + +### 15 — Issue #123 × *Recursion* + +**EN** The past did not return; it had never finished leaving the room. + +**中文** 过去并未归来;它只是从未彻底离开这间屋子。 + +### 16 — Issue #15 × *Death’s End* + +**EN** The field divided left from right, but the old silver coat remembered the same wind. + +**中文** 麦田分出了左与右,那件旧银衣却记得同一阵风。 + +### 17 — Hydrodynamics × *Dune* + +**EN** No grain of sand knows the river, yet the desert flows. + +**中文** 没有一粒沙知道河流,沙漠却依然流动。 + +### 18 — Issue #158 × *The Dark Forest* + +**EN** At first every tree stood alone; then the whole forest chose a direction. + +**中文** 起初每棵树都独自站立,后来整座森林选择了同一个方向。 + +### 19 — Issue #128 × *Foundation* + +**EN** A future becomes trustworthy not when we prefer it, but when we can still bound its error. + +**中文** 一个未来值得信任,不因我们偏爱它,而因它的误差仍可被界定。 + +### 20 — Issue #133 × *The Last Question* + +**EN** The last machine did not print an answer; it released a blank card into the morning. + +**中文** 最后的机器没有印出答案,只把一张空白卡片送进清晨。 + diff --git a/docs/showcase/ranger-archive/PROMPTS-v2.md b/docs/showcase/ranger-archive/PROMPTS-v2.md new file mode 100644 index 000000000..c22db7478 --- /dev/null +++ b/docs/showcase/ranger-archive/PROMPTS-v2.md @@ -0,0 +1,85 @@ +# Lost Science-Fiction Covers V2 — GPT Image Prompt Set + +## Shared production constraints + +Every image is an original panoramic cover illustration for a research pull request, approximately 2.5:1. It belongs to a different publishing era and must not inherit the visual style of the other covers. Show human presence through worn places and used objects rather than a posed protagonist. Make the referenced science-fiction work recognizable through an iconic ordinary setting or prop, but do not reproduce an existing cover, film frame, actor likeness, character costume, franchise logo, or proprietary spacecraft. The scientific subject appears as one physical anomaly inside an otherwise believable scene. No title, caption, readable prose, watermark, UI overlay, decorative outer-space background, generic neon network, recurring probe, or unnecessary glowing ring. Leave some visually quiet space suitable for later PR typography. + +## 01 — The First Ratio + +Use case: stylized-concept. Asset: wide PR cover. Create an original early-1970s large-format color landscape photograph inspired thematically by *2001: A Space Odyssey*. At cold dawn, a severe matte-black rectangular slab stands in a dew-covered green grassland. On one side, modest wooden survey stakes mark a triangular planting plot; on the other, stakes mark a honeycomb planting plot. A weathered canvas survey bag, folding brass ruler, muddy notebook, and thermos lie abandoned in the foreground, showing that a field researcher just stepped away. The brass rods quietly establish a √5-like proportional relation without displaying a formula. Pale sunrise, fog, real grass and mud, restrained composition, tactile analog film grain. No people, apes, spacecraft, desert, star field, text, logo, neon geometry, or copied movie frame. + +## 02 — The Planet Sings + +Use case: historical-scene. Asset: wide PR cover. Create an original 1960s Chinese gouache science-poster painting, thematically recalling *The Wandering Earth*. A snowbound northern school courtyard at blue winter morning: brick classrooms, bicycle tracks, bare poplar trees, and enormous planetary engine towers barely visible beyond ordinary rooftops. On a classroom windowsill sits a chipped enamel lunch bowl filled with water, a red wool scarf, chalk, and a small crystal radio. The water, the glass window, and distant engine exhaust all share the same subtle physical vibration pattern, expressing electron–phonon coupling from tabletop to planet. Opaque brushwork, faded revolutionary-era printing pigments, warm windows against cold snow. No slogan, legible characters, people, logos, deep-space view, circles, or generic luminous network. + +## 03 — The Untested Machine + +Use case: photorealistic-natural. Asset: wide PR cover. Create an original faded instant-film field photograph inspired by *The Martian*. Inside a cramped Mars greenhouse/workshop, red dust presses against a scratched window. A wooden bench holds sprouting potato trays, silver duct tape, a wrench, stained gloves, sample jars, and a newly repaired compact tensor-computing device with mismatched rust-colored parts. A row of simple physical test objects and indicator lamps shows repeated verification; one lamp is green, several earlier attempts are crossed out only as non-readable marks. Harsh practical habitat light, imperfect instant-film exposure, fingerprints, condensation, human improvisation everywhere. No astronaut, actor likeness, readable checklist text, title, logo, star field, abstract rings, or polished futuristic laboratory. + +## 04 — The Answer Came First + +Use case: illustration-story. Asset: wide PR cover. Create an original eccentric 1980s British comic-paperback illustration inspired by *The Hitchhiker’s Guide to the Galaxy*. Keep the composition on a rain-spattered roadside café table and the anonymous wet road beyond it; do not include a shopfront, signboard, menu, poster, advertising, or wall writing. The table holds a striped towel, chipped teacup, half-eaten toast, cheap ballpoint pen, and a wildly overfolded road atlas. A crumpled receipt shows only the large number 42. Hand-drawn arrows and absurd detours run backward from 42 through the atlas toward one completely blank square. Mustard yellow, petrol blue, tomato red, uneven ink outlines, dry visual humor, cheap offset-print texture. No other readable text, question mark, space scene, glowing portal, people, copyrighted logo, or exact cover imitation. + +## 05 — A Language That Refuses Falsehood + +Use case: stylized-concept. Asset: wide PR cover. Create an original 1950s three-color silkscreen book illustration inspired by *I, Robot*. An orderly drafting office after hours: green metal desk, manual typewriter, pencil sharpener, eraser crumbs, coffee stain, and an articulated brushed-metal hand resting beside three punched rule cards. A branching paper syntax diagram physically threads through the three cards; one well-formed branch reaches a finished stack, while a malformed red branch is mechanically clipped and dropped into a wastebasket. Flat cream, black, institutional green and safety red inks, slight registration errors, mid-century graphic economy. No humanoid robot face, actor, readable sentences, neon ring, star field, title, logo, or photoreal 3D render. + +## 06 — The Phase Remembers + +Use case: stylized-concept. Asset: wide PR cover. Create an original translucent watercolor and ink-wash illustration inspired by *Arrival*. A foggy Montana-like meadow outside a temporary language classroom at dawn; folding chair, rubber gloves, damp notebook, thermos, and a large rain-streaked glass panel show recent human work. On the glass is one imperfect black circular ink gesture, the only dominant circle in the series. Its beginning and end nearly meet, but a small iridescent offset at the seam records geometric phase after a closed journey. Soft gray fog, paper fibers, watery edges, muted violet and moss green. No aliens, people, helicopters, readable writing, repeated rings, outer space, logo, or copied film composition. + +## 07 — The Shape of State Space + +Use case: historical-scene. Asset: wide PR cover. Create an original 1880s hand-colored copper engraving inspired by *Flatland*. A Victorian study table holds spectacles, ruler, sealed letter, teacup, crumbs, and a checked linen tablecloth printed with tiny two-dimensional geometric inhabitants. At the center, the supposedly flat engraved grid physically lifts, folds, and becomes a gently curved surface while the ink figures remain trapped in two dimensions. Cross-hatched copper lines, ivory paper, restrained hand-applied vermilion, indigo, and sage pigments, antique book-plate imperfections. No circles as focal form, no neon, no computer graphics, no stars, no readable text, no modern objects. + +## 08 — The Horizon of Proof + +Use case: historical-scene. Asset: wide PR cover. Create an original medieval illuminated-manuscript and egg-tempera scene inspired by *Anathem*. A sunlit stone cloister contains an open proof book, wax tablet, brass divider, worn sandals, and an extinguished candle, but no visible monk. Two massive columns cast precisely separated shadows across the floor; between the shadow edges lies one narrow band of gold leaf representing a certified energy value bounded from below and above. Mineral blue, ochre, chalk white, cracked tempera, handmade vellum texture. No futuristic towers, outer space, equations, readable script, circles, neon glow, character portrait, or copied cover. + +## 09 — Rebuild the Tool Between Stars + +Use case: stylized-concept. Asset: wide PR cover. Create an original 1970s technical-airbrush paperback illustration inspired by *Project Hail Mary*. A cluttered interstellar laboratory workbench is divided by two visibly different environmental zones: one human side with coffee pouch, scorched circuit board, screwdriver and taped notes; one alien-compatible side with faceted transparent structural material and a handmade musical chime object. Across the bench, a fragile pale circuit has been rebuilt into a rugged rust-colored computational core. Warm amber practical light and cool ammonia-blue enclosure, analog airbrush gradients, slightly worn mass-market print. No astronaut, alien character, franchise ship, readable text, star-field background, central ring, logo, or glossy modern concept art. + +## 10 — The Efficient Path + +Use case: stylized-concept. Asset: wide PR cover. Create an original 1980s arcade-box screenprint inspired by *Ender’s Game*. An empty military-school dormitory desk after lights-out holds scuffed shoes, a juice carton, pencil stubs, and a small handheld tactical board with many cheap wooden fleet pieces. Most attempted routes are marked by tangled white grease-pencil strokes, but one extremely short gold move solves the formation using far fewer pieces. Electric cobalt, magenta, black, and gold, coarse halftone dots and dramatic diagonal composition. No child, actor, battle-room sphere, readable interface, logo, outer-space fleet, glowing ring, or contemporary 3D rendering. + +## 11 — What the Cycle Carries + +Use case: stylized-concept. Asset: wide PR cover. Create an original industrial photo-collage with shipping-label textures inspired by *The Expanse*. A worn asteroid cargo dock contains steel loading bays, scuffed floor paint, grease pencil, work gloves, ration tin, hand trolley, and a physical rectangular conveyor system. The conveyor completes a mechanical cycle back to its starting configuration, yet one sealed parcel has advanced exactly one numbered bay; bay numbers may be simple digits, with no other readable text. Rust, sodium-vapor orange, dirty ice blue, torn-paper collage edges. No circular gate, franchise logo, people, spaceship glamour shot, neon network, outer-space background, or copied production design. + +## 12 — Crossing the State-Space Horizon + +Use case: photorealistic-natural. Asset: wide PR cover. Create an original naturalistic 35mm rural cinema still thematically inspired by *Interstellar*. A lush green cornfield and grassy verge lead to a weathered white farmhouse under a huge late-afternoon sky. On the porch are a child’s handmade space model, dusty wristwatch, chalk, and a freestanding blackboard covered only in an abstract determinant-grid pattern, not readable equations. The grid’s straight rows subtly bend toward one small gravitational distortion high in the clouds. Warm sun on grass, approaching storm, real film grain, emotionally grounded American farm life. No people, actor likeness, branded tractor, spacecraft, black-hole close-up, glowing orbit, text, or copied movie frame. + +## 13 — The Path Without a Sign + +Use case: photorealistic-natural. Asset: wide PR cover. Create an original gritty 1970s Eastern European color documentary photograph inspired by *Roadside Picnic*. An overgrown rail yard after rain: rusted tracks, concrete pylons, weeds, muddy water, abandoned sandwich tin, cheap raincoat, survey flags, and boot prints. Along false routes, nuts and bolts float at impossible heights and shadows point the wrong way. One utterly ordinary trail of wet boot prints passes safely through the Zone without interference. Desaturated green, brown and oxidized orange, scratched film, low cloudy light. No stalker character, weapon, text, supernatural portal, circles, neon paths, star field, or polished game concept art. + +## 14 — The Ocean Answers Back + +Use case: historical-scene. Asset: wide PR cover. Create an original 1960s Polish art-film still inspired by *Solaris*. A quiet modernist dacha room contains an empty wooden chair, half-cut apple browning on a plate, rumpled blanket, analog oscilloscope, open drawer, and condensation on a wide window overlooking a strange silver ocean. The ocean surface repeats the oscilloscope waveform a moment later, as if the environment has learned the room’s critical fluctuation. Muted olive, tobacco brown, pearl gray, natural window light, subtle film dust and psychological stillness. No person, face, astronaut, spaceship, giant glowing ocean ring, text, logo, or copied Tarkovsky frame. + +## 15 — Memory Inside the Loop + +Use case: photorealistic-natural. Asset: wide PR cover. Create an original 1990s family-photo contact sheet with VHS color bleed, inspired by *Recursion*. Four adjacent rectangular photographs show the same modest kitchen at four dates: floral tablecloth, birthday candle, cassette recorder, refrigerator magnets without readable words, a child’s cup, and an empty chair. Objects from the earliest photograph physically leak into later frames—an old candle shadow appears before the candle, spilled milk reverses direction, the recorder’s red light persists—making memory visibly non-Markovian without using a circular loop. Flash photography, faded cyan and warm skinless domestic colors, date-stamp shapes but no readable date. No people, faces, text, spirals, cosmic background, or sleek sci-fi interface. + +## 16 — Gravity Chooses a Hand + +Use case: illustration-story. Asset: wide PR cover. Create an original contemporary Chinese literary-cover painting inspired by *Death’s End*. A sunlit golden wheat field grows from pure black crystalline soil inside an immense pale artificial habitat. An old silver reflective jacket and straw hat hang on a simple scarecrow at the left edge, suggesting a long-absent traveler. The field is split by one subtle vertical crease in the air: every wheat stalk on the left leans in straight parallel diagonal rows toward the far left, while every stalk on the right leans in straight parallel diagonal rows toward the far right, expressing opposite gravitational handedness. No spirals, circles, vortices, rings, braided coils, curved crop patterns, or concentric formations. Lyrical oil-and-gouache texture, warm harvest gold, ink-black crystalline foreground, pale silver accents, quiet sorrow. No visible character, spaceship, text, logo, star field, or recreation of an existing cover. + +## 17 — From Grains to Currents + +Use case: stylized-concept. Asset: wide PR cover. Create an original Japanese woodblock print with mineral pigments, inspired by *Dune*. A desert survey camp at high morning contains a wrapped water flask, folded indigo cloth, measuring poles, shaded notebook, and footprints being erased by wind. Close foreground grains are individually carved as dots; farther away they merge into sweeping river-like streamlines and one breaking fluid vortex around a rock, showing hydrodynamics emerging from particles. Sand gold, indigo, persimmon red and unprinted paper, visible wood grain, bold asymmetry. No person, worm, franchise costume, twin moons, text, glowing rings, photoreal 3D, or copied film imagery. + +## 18 — When the Forest Orders Itself + +Use case: stylized-concept. Asset: wide PR cover. Create an original rough charcoal drawing combined with black linocut, inspired by *The Dark Forest*. A snowy forest clearing at dusk contains a cold camp stove, closed radio notebook, thermos cup, one broken antenna, and rows of distant trunks. On the left, tree marks and falling snow are locally random; across the wide composition, almost imperceptibly, trunks, branches, animal tracks and radio scratches align into long-range order despite the cold silence. Carbon black, paper white, a trace of iron red, aggressive carved texture. No hunter, gun, people, planets, neon network, circles, glowing stars, readable text, or literal book-cover copy. + +## 19 — The Future Inside a Bound + +Use case: stylized-concept. Asset: wide PR cover. Create an original 1950s modernist paper collage inspired by *Foundation*. A vast public reading room beneath a layered city dome contains a green librarian’s lamp, empty spectacles, punched cards, paper timelines, scissors and archival boxes. Numerous paper strips branch toward different futures. A transparent amber drafting ruler forms a strict error boundary: only trajectories inside it remain crisp; strips outside become torn, misregistered, and unreliable. Cut-paper geometry, Bauhaus-era cream, burgundy, teal and amber, tactile glue shadows. No people, actor likeness, galaxy, spacecraft, glowing vault, circles as focus, readable text, logo, or glossy digital render. + +## 20 — The Seed of the Next Question + +Use case: stylized-concept. Asset: wide PR cover. Create an original 1960s two-color risograph educational-magazine illustration inspired by *The Last Question*. An empty classroom at sunrise contains chalk dust, wooden desks, a coat left on a chair, pencil shavings, and stacks of question cards—some stamped only with abstract checks or crosses, no readable words—feeding into a large electromechanical problem press assembled from school duplicators and early computers. After multiple physical validation gates, the machine releases one pristine blank card into the first ray of daylight: the seed of the next problem. Cobalt blue, fluorescent orange and warm paper, risograph misregistration, hopeful handmade machinery. No cosmic computer face, people, readable title, outer space, glowing ring, logo, or modern UI. diff --git a/docs/showcase/ranger-archive/PROMPTS.md b/docs/showcase/ranger-archive/PROMPTS.md new file mode 100644 index 000000000..8ca72403f --- /dev/null +++ b/docs/showcase/ranger-archive/PROMPTS.md @@ -0,0 +1,68 @@ +# Ranger Archive — GPT Image Prompt Set + +## Shared prompt + +Create an original cinematic hard-science-fiction panoramic banner, approximately 2.5:1. Use deep-space black, restrained luminous accents, crisp scientific structures, atmospheric depth, and one — exactly one — tiny diamond-shaped Ranger probe as the recurring observer. The probe must remain secondary to the scientific phenomenon. No text, captions, equations used as labels, logos, watermarks, recognizable franchise characters, copied props, or direct recreation of a copyrighted film frame. The referenced science-fiction work is a thematic lens only. The image must tell one clear story at thumbnail size and reserve some quiet negative space for optional PR typography. + +## Mission prompts + +1. **The First Ratio — Issue #148 / _2001: A Space Odyssey_** + A severe black monolith of seemingly absolute proportion stands between luminous triangular and honeycomb lattices; their critical scales converge toward a mysterious √5-like relation while the probe measures the gap. + +2. **When the Planet Sings — Issue #35 / _The Wandering Earth_** + Planet-scale lattice engines couple electronic light to vast phonon ripples beneath an icebound world; the probe listens where microscopic vibration becomes planetary motion. + +3. **The Untested Machine — Issue #114 / _The Martian_** + An isolated rust-red verification habitat surrounds a newly assembled tensor machine; test beams, reference signals, and failed components reveal survival through methodical validation. + +4. **The Answer Came First — Issue #71 / _The Hitchhiker’s Guide to the Galaxy_** + A luminous answer symbol at the far end of a comic-cosmic reverse-search star map; branching arithmetic routes travel backward toward the missing question. + +5. **A Language That Refuses Falsehood — Issue #82 / _I, Robot_** + A crystalline syntax tree is guarded by three concentric rule rings; valid tensor operations pass as coherent light while an invalid branch is rejected outside the boundary. + +6. **The Phase Remembers — Issue #73 / _Arrival_** + A closed circular quantum glyph formed from overlapping states; after one orbit, a small luminous seam records geometric phase while local coordinates return unchanged. + +7. **The Shape of State Space — Quantum Geometry / _Flatland_** + A perfectly flat grid curls upward into a luminous curved manifold; inhabitants remain planar while the probe crosses the boundary and sees distance become geometry. + +8. **The Horizon of Proof — Issue #230 / _Anathem_** + Two austere proof towers mark certified lower and upper energy bounds; between them lies a narrow radiant horizon containing the exact Bethe-ansatz value. + +9. **Rebuild the Tool Between Stars — Issue #115 / _Project Hail Mary_** + Inside a remote interstellar laboratory, a fragile circuit prototype is reconstructed into a rugged rust-colored computational core capable of surviving alone. + +10. **The Efficient Path — Issue #33 / _Ender’s Game_** + A tactical tensor-network arena presents countless costly routes; one precise gold trajectory achieves the target state using radically fewer computational resources. + +11. **What the Cycle Carries — Issue #36 / _The Expanse_** + An original ring-shaped transport gate encloses a modulated lattice; a particle packet completes a parameter cycle and emerges displaced by one quantized cell. + +12. **Crossing the State-Space Horizon — Issue #129 / _Interstellar_** + A black-hole-like Hamiltonian bends an immense constellation of many-body determinants; a narrow calculation beam charts the survivable path through exact diagonalization. + +13. **The Path Without a Sign — Issue #121 / _Roadside Picnic_** + A forbidden quantum zone is filled with seductive but unstable routes; one machine-verified corridor remains free of destructive sign interference. + +14. **The Ocean Answers Back — Issue #122 / _Solaris_** + A dark responsive ocean surrounds an open-system apparatus; critical fluctuations return the observer’s signal in an altered form, blurring system and environment. + +15. **Memory Inside the Loop — Issue #123 / _Recursion_** + Nested luminous Floquet loops retain traces of previous cycles; a delayed echo re-enters the present and changes the next orbit. + +16. **Gravity Chooses a Hand — Issue #15 / _Death’s End_** + Cyan and magenta chiral gravitational waves pass through a folding dimensional membrane, entangling with a neural quantum-state web. + +17. **From Grains to Currents — Hydrodynamics / _Dune_** + A desert of individually visible grains transitions into immense coherent streamlines and vortices; the probe witnesses microscopic collisions becoming fluid law. + +18. **When the Forest Orders Itself — Issue #158 / _The Dark Forest_** + Weak isolated correlations in a nearly black forest of points cross a frontier and become a globally aligned luminous network, challenging the expected loss of order. + +19. **The Future Inside a Bound — Issue #128 / _Foundation_** + Many simulated futures branch beneath nested error vaults; trajectories outside the strict Trotter bound dissolve while certified branches remain sharp. + +20. **The Seed of the Next Question — Issue #133 / _The Last Question_** + Nineteen colored mission paths converge into a cosmic archive; after passing transparent human-approval and verification gates, one new golden problem seed exits toward first light. + diff --git a/docs/showcase/ranger-archive/README.md b/docs/showcase/ranger-archive/README.md new file mode 100644 index 000000000..91c057a9c --- /dev/null +++ b/docs/showcase/ranger-archive/README.md @@ -0,0 +1,55 @@ +# Lost Science-Fiction Covers — Ranger Archive V2 + +这是为 20 个研究项目重新设计的原创科幻横幅。V2 不再追求统一的深空概念图,而把每个项目设计成一本来自不同年代、不同出版文化的“遗失科幻书”: + +![V2 的 20 个任务 4×5 总览](contact-sheet-v2.png) + +旧版图片仍保存在 [`assets/missions`](assets/missions),[旧版总览](contact-sheet.png)也保留作比较;V2 图片位于 [`assets/missions-v2`](assets/missions-v2)。 + +## V2 视觉原则 + +- 画幅:`1983 × 793`,约 `2.5:1`,适合 PR 顶部横幅。 +- 二十种出版媒介:风景摄影、水粉科普画、拍立得、漫画、丝网印刷、水彩、铜版画、蛋彩画、技术喷绘、街机印刷、拼贴、35mm 胶片、纪实摄影、艺术电影、家庭相册、文学绘画、木版画、炭笔版画、现代主义拼贴和孔版印刷。 +- 人不一定出镜,但必须留下使用痕迹:围巾、搪瓷盆、土豆、胶带、杯子、鞋、手套、相册、旧外套、笔记本和灯。 +- 科幻作品通过一眼可辨的普通环境或物件进入画面,不复刻演员、角色造型、飞船或电影镜头。 +- 科学主题表现为现实里的一处异常,不再默认画成发光圆环、轨道和网络。 +- V2 完全取消了统一的 Ranger 探测器。 +- 图中不放标题、公式说明、品牌、作品角色或标识;文字由 PR 正文承载。 +- 科幻作品只作为叙事隐喻,不复刻受版权保护的具体镜头、人物或美术资产。 + +## 二十个任务 + +| # | 研究项目 | 科幻作品镜头 | 画面叙事 | 文件 | +|---:|---|---|---|---| +| 01 | Issue #148:临界点比值与 √5 | 《2001:太空漫游》 | 黎明草地上的黑色石碑、测量包和两种种植网格 | [`01-issue-148-first-ratio-v2.png`](assets/missions-v2/01-issue-148-first-ratio-v2.png) | +| 02 | Issue #35:电声相互作用 | 《流浪地球》 | 雪窗、搪瓷盆和老收音机与远处行星发动机共同振动 | [`02-issue-35-planet-sings-v2.png`](assets/missions-v2/02-issue-35-planet-sings-v2.png) | +| 03 | Issue #114:Rust 张量库验证 | 《火星救援》 | 火星温室里的土豆、胶带和逐项受测的修补机器 | [`03-issue-114-untested-machine-v2.png`](assets/missions-v2/03-issue-114-untested-machine-v2.png) | +| 04 | Issue #71:隐函数恢复 | 《银河系漫游指南》 | 雨中公路餐桌上的毛巾、茶和从 42 倒推的破旧地图 | [`04-issue-71-answer-first-v2.png`](assets/missions-v2/04-issue-71-answer-first-v2.png) | +| 05 | Issue #82:Certified tensor DSL | 《我,机器人》 | 三张规则卡让正确纸带通过,把错误分支剪进废纸篓 | [`05-issue-82-language-of-rules-v2.png`](assets/missions-v2/05-issue-82-language-of-rules-v2.png) | +| 06 | Issue #73:2D-TFIM 几何相位 | 《降临》 | 雾地语言站的湿玻璃上,墨迹接缝留下微小相位偏移 | [`06-issue-73-phase-remembers-v2.png`](assets/missions-v2/06-issue-73-phase-remembers-v2.png) | +| 07 | Quantum Geometry | 《平面国》 | 维多利亚书桌上的格纹桌布从平面隆起为曲面 | [`07-quantum-geometry-curved-state-space-v2.png`](assets/missions-v2/07-quantum-geometry-curved-state-space-v2.png) | +| 08 | Issue #230:Bethe ansatz 能量密度界 | 《教典》 | 修道院两根柱子的阴影夹住唯一的金色证明区间 | [`08-issue-230-proof-horizon-v2.png`](assets/missions-v2/08-issue-230-proof-horizon-v2.png) | +| 09 | Issue #115:Occam Circuit Rust 移植 | 《挽救计划》 | 人类咖啡杯与异星材料之间,一台脆弱工具被重新造好 | [`09-issue-115-rebuild-the-tool-v2.png`](assets/missions-v2/09-issue-115-rebuild-the-tool-v2.png) | +| 10 | Issue #33:Extreme-efficiency VQE | 《安德的游戏》 | 宿舍战术棋盘上,一步金色短解穿过无数失败涂痕 | [`10-issue-33-efficient-path-v2.png`](assets/missions-v2/10-issue-33-efficient-path-v2.png) | +| 11 | Issue #36:Interacting Thouless pump | 《苍穹浩瀚》 | 小行星货运码头完成一个机械周期,货箱却前进一格 | [`11-issue-36-cycle-transports-v2.png`](assets/missions-v2/11-issue-36-cycle-transports-v2.png) | +| 12 | Issue #129:ED/FCI workbench | 《星际穿越》 | 玉米地、旧农舍、儿童模型和弯向天空异常的黑板网格 | [`12-issue-129-state-space-crossing-v2.png`](assets/missions-v2/12-issue-129-state-space-crossing-v2.png) | +| 13 | Issue #121:Sign-problem-free hunter | 《路边野餐》 | 雨后禁区中零件悬浮,只有普通泥脚印仍然安全 | [`13-issue-121-safe-path-v2.png`](assets/missions-v2/13-issue-121-safe-path-v2.png) | +| 14 | Issue #122:开放量子物质临界性 | 《索拉里斯星》 | 空椅、半只苹果和窗外重复示波器波形的海 | [`14-issue-122-ocean-answers-v2.png`](assets/missions-v2/14-issue-122-ocean-answers-v2.png) | +| 15 | Issue #123:非 Markovian Floquet 系统 | 《递归》 | 四张家庭相片中,旧日物件和影子侵入后来的时刻 | [`15-issue-123-memory-loops-v2.png`](assets/missions-v2/15-issue-123-memory-loops-v2.png) | +| 16 | Issue #15:Chiral graviton | 《死神永生》 | 黑色晶体土壤上的麦田向相反手性倒伏,留下银色外套 | [`16-issue-15-chiral-gravity-v2.png`](assets/missions-v2/16-issue-15-chiral-gravity-v2.png) | +| 17 | Hydrodynamics | 《沙丘》 | 木版沙漠中,单颗沙粒渐渐汇成宏观流线与涡流 | [`17-hydrodynamics-emergent-flow-v2.png`](assets/missions-v2/17-hydrodynamics-emergent-flow-v2.png) | +| 18 | Issue #158:长程 Mermin–Wagner | 《黑暗森林》 | 黑白雪林从随机枝干逐渐排列成长程秩序 | [`18-issue-158-long-range-order-v2.png`](assets/missions-v2/18-issue-158-long-range-order-v2.png) | +| 19 | Issue #128:Trotter 严格误差界 | 《基地》 | 图书馆纸带的众多未来被一把透明误差尺严格裁定 | [`19-issue-128-bounded-future-v2.png`](assets/missions-v2/19-issue-128-bounded-future-v2.png) | +| 20 | Issue #133:The Problem Factory | 《最后的问题》 | 清晨教室里的问题印刷机筛选卡片,送出下一张空白问题 | [`20-issue-133-question-seed-v2.png`](assets/missions-v2/20-issue-133-question-seed-v2.png) | + +## PR 中的推荐用法 + +20 组可直接配图使用的英文+中文原创题记见 [`CAPTIONS-v2.md`](CAPTIONS-v2.md)。 + +每个项目只使用自己的横幅,并在图片下方放置三层文字: + +1. **作品短引文**:不超过一句,并在发布前核对版本与译文版权; +2. **原创题记**:把科幻母题转译为这个项目的认识论动作; +3. **项目状态句**:只陈述本 PR 真正完成、证明或尚未跨过的验收门。 + +整组总结 PR 仍可按照五幕顺序排列,每幕先给一段短序,再展示四张项目卡。V2 的系列感来自“二十本遗失的科幻书”,不再依赖同一个角色、配色或构图。 diff --git a/docs/showcase/ranger-archive/assets/missions-v2/01-issue-148-first-ratio-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/01-issue-148-first-ratio-v2.png new file mode 100644 index 000000000..30accc87d Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/01-issue-148-first-ratio-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/02-issue-35-planet-sings-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/02-issue-35-planet-sings-v2.png new file mode 100644 index 000000000..9226664b9 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/02-issue-35-planet-sings-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/03-issue-114-untested-machine-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/03-issue-114-untested-machine-v2.png new file mode 100644 index 000000000..b9dac679c Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/03-issue-114-untested-machine-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/04-issue-71-answer-first-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/04-issue-71-answer-first-v2.png new file mode 100644 index 000000000..6e7c01f3f Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/04-issue-71-answer-first-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/05-issue-82-language-of-rules-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/05-issue-82-language-of-rules-v2.png new file mode 100644 index 000000000..d2786238c Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/05-issue-82-language-of-rules-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/06-issue-73-phase-remembers-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/06-issue-73-phase-remembers-v2.png new file mode 100644 index 000000000..cd2281183 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/06-issue-73-phase-remembers-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/07-quantum-geometry-curved-state-space-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/07-quantum-geometry-curved-state-space-v2.png new file mode 100644 index 000000000..e7343447d Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/07-quantum-geometry-curved-state-space-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/08-issue-230-proof-horizon-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/08-issue-230-proof-horizon-v2.png new file mode 100644 index 000000000..bdd9d168d Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/08-issue-230-proof-horizon-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/09-issue-115-rebuild-the-tool-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/09-issue-115-rebuild-the-tool-v2.png new file mode 100644 index 000000000..002274fe7 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/09-issue-115-rebuild-the-tool-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/10-issue-33-efficient-path-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/10-issue-33-efficient-path-v2.png new file mode 100644 index 000000000..69a91df4b Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/10-issue-33-efficient-path-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/11-issue-36-cycle-transports-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/11-issue-36-cycle-transports-v2.png new file mode 100644 index 000000000..c241b844e Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/11-issue-36-cycle-transports-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/12-issue-129-state-space-crossing-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/12-issue-129-state-space-crossing-v2.png new file mode 100644 index 000000000..0711e0849 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/12-issue-129-state-space-crossing-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/13-issue-121-safe-path-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/13-issue-121-safe-path-v2.png new file mode 100644 index 000000000..f10784473 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/13-issue-121-safe-path-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/14-issue-122-ocean-answers-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/14-issue-122-ocean-answers-v2.png new file mode 100644 index 000000000..568d23194 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/14-issue-122-ocean-answers-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/15-issue-123-memory-loops-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/15-issue-123-memory-loops-v2.png new file mode 100644 index 000000000..9fbbc2d00 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/15-issue-123-memory-loops-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/16-issue-15-chiral-gravity-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/16-issue-15-chiral-gravity-v2.png new file mode 100644 index 000000000..fe19dde2a Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/16-issue-15-chiral-gravity-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/17-hydrodynamics-emergent-flow-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/17-hydrodynamics-emergent-flow-v2.png new file mode 100644 index 000000000..25aa5f66c Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/17-hydrodynamics-emergent-flow-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/18-issue-158-long-range-order-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/18-issue-158-long-range-order-v2.png new file mode 100644 index 000000000..fa9864693 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/18-issue-158-long-range-order-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/19-issue-128-bounded-future-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/19-issue-128-bounded-future-v2.png new file mode 100644 index 000000000..155baf598 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/19-issue-128-bounded-future-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions-v2/20-issue-133-question-seed-v2.png b/docs/showcase/ranger-archive/assets/missions-v2/20-issue-133-question-seed-v2.png new file mode 100644 index 000000000..36cd6900e Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions-v2/20-issue-133-question-seed-v2.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/01-issue-148-first-ratio.png b/docs/showcase/ranger-archive/assets/missions/01-issue-148-first-ratio.png new file mode 100644 index 000000000..b5fe167ca Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/01-issue-148-first-ratio.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/02-issue-35-planet-sings.png b/docs/showcase/ranger-archive/assets/missions/02-issue-35-planet-sings.png new file mode 100644 index 000000000..844a33b21 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/02-issue-35-planet-sings.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/03-issue-114-untested-machine.png b/docs/showcase/ranger-archive/assets/missions/03-issue-114-untested-machine.png new file mode 100644 index 000000000..133f74a2d Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/03-issue-114-untested-machine.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/04-issue-71-answer-first.png b/docs/showcase/ranger-archive/assets/missions/04-issue-71-answer-first.png new file mode 100644 index 000000000..83cf31fc8 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/04-issue-71-answer-first.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/05-issue-82-language-of-rules.png b/docs/showcase/ranger-archive/assets/missions/05-issue-82-language-of-rules.png new file mode 100644 index 000000000..987fdb3d5 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/05-issue-82-language-of-rules.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/06-issue-73-phase-remembers.png b/docs/showcase/ranger-archive/assets/missions/06-issue-73-phase-remembers.png new file mode 100644 index 000000000..d5a6d47f2 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/06-issue-73-phase-remembers.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/07-quantum-geometry-curved-state-space.png b/docs/showcase/ranger-archive/assets/missions/07-quantum-geometry-curved-state-space.png new file mode 100644 index 000000000..d448fa768 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/07-quantum-geometry-curved-state-space.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/08-issue-230-proof-horizon.png b/docs/showcase/ranger-archive/assets/missions/08-issue-230-proof-horizon.png new file mode 100644 index 000000000..afc42dbb4 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/08-issue-230-proof-horizon.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/09-issue-115-rebuild-the-tool.png b/docs/showcase/ranger-archive/assets/missions/09-issue-115-rebuild-the-tool.png new file mode 100644 index 000000000..6c0e0fc03 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/09-issue-115-rebuild-the-tool.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/10-issue-33-efficient-path.png b/docs/showcase/ranger-archive/assets/missions/10-issue-33-efficient-path.png new file mode 100644 index 000000000..ce924383d Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/10-issue-33-efficient-path.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/11-issue-36-cycle-transports.png b/docs/showcase/ranger-archive/assets/missions/11-issue-36-cycle-transports.png new file mode 100644 index 000000000..3a2d8d4cf Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/11-issue-36-cycle-transports.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/12-issue-129-state-space-crossing.png b/docs/showcase/ranger-archive/assets/missions/12-issue-129-state-space-crossing.png new file mode 100644 index 000000000..1301beeef Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/12-issue-129-state-space-crossing.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/13-issue-121-safe-path.png b/docs/showcase/ranger-archive/assets/missions/13-issue-121-safe-path.png new file mode 100644 index 000000000..0239c81fa Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/13-issue-121-safe-path.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/14-issue-122-ocean-answers.png b/docs/showcase/ranger-archive/assets/missions/14-issue-122-ocean-answers.png new file mode 100644 index 000000000..29fb04f60 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/14-issue-122-ocean-answers.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/15-issue-123-memory-loops.png b/docs/showcase/ranger-archive/assets/missions/15-issue-123-memory-loops.png new file mode 100644 index 000000000..1b8c7d8af Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/15-issue-123-memory-loops.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/16-issue-15-chiral-gravity.png b/docs/showcase/ranger-archive/assets/missions/16-issue-15-chiral-gravity.png new file mode 100644 index 000000000..39a8546ff Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/16-issue-15-chiral-gravity.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/17-hydrodynamics-emergent-flow.png b/docs/showcase/ranger-archive/assets/missions/17-hydrodynamics-emergent-flow.png new file mode 100644 index 000000000..7c6d94e76 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/17-hydrodynamics-emergent-flow.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/18-issue-158-long-range-order.png b/docs/showcase/ranger-archive/assets/missions/18-issue-158-long-range-order.png new file mode 100644 index 000000000..f6497830d Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/18-issue-158-long-range-order.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/19-issue-128-bounded-future.png b/docs/showcase/ranger-archive/assets/missions/19-issue-128-bounded-future.png new file mode 100644 index 000000000..f64c3c953 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/19-issue-128-bounded-future.png differ diff --git a/docs/showcase/ranger-archive/assets/missions/20-issue-133-question-seed.png b/docs/showcase/ranger-archive/assets/missions/20-issue-133-question-seed.png new file mode 100644 index 000000000..74ec4cbf6 Binary files /dev/null and b/docs/showcase/ranger-archive/assets/missions/20-issue-133-question-seed.png differ diff --git a/docs/showcase/ranger-archive/contact-sheet-v2.png b/docs/showcase/ranger-archive/contact-sheet-v2.png new file mode 100644 index 000000000..6fb38d1a3 Binary files /dev/null and b/docs/showcase/ranger-archive/contact-sheet-v2.png differ diff --git a/docs/showcase/ranger-archive/contact-sheet.png b/docs/showcase/ranger-archive/contact-sheet.png new file mode 100644 index 000000000..32eb47169 Binary files /dev/null and b/docs/showcase/ranger-archive/contact-sheet.png differ diff --git a/tracks/agent-kb/solutions/WangTheoPhys/.gitignore b/tracks/agent-kb/solutions/WangTheoPhys/.gitignore new file mode 100644 index 000000000..a160f494b --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/.gitignore @@ -0,0 +1,3 @@ +__pycache__/ +*.py[cod] +.pytest_cache/ diff --git a/tracks/agent-kb/solutions/WangTheoPhys/README.md b/tracks/agent-kb/solutions/WangTheoPhys/README.md new file mode 100644 index 000000000..d4eef8a6f --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/README.md @@ -0,0 +1,315 @@ +# WangTheoPhys: a fail-closed tensor-network research gate + +## Team + +| | | +|---|---| +| **Team name** | WangTheoPhys | +| **Members** | Junkai Wang, WangTheoPhys@outlook.com | + +## Challenge + +| Row | | +|---|---| +| **Challenge** | Build an agent for tensor-network research: a research assistant that can mine tensor-network literature, maintain a grounded knowledge base, propose executable research tasks, and support solution workflows. | +| **Catalog issue** | Addresses [#133 — The problem factory](https://github.com/QuantumBFS/quantum.harness/issues/133), released by Jin-Guo Liu. | +| **Track** | `agent-kb`, chosen because the issue's `Method` field is `Other` and this is an AI research-agent / knowledge-base contribution. | + +This directory is a deliberately small public review surface for the larger +TN-Agent design. It preregisters scientific intent, binds an exact executable +route, evaluates normalized evidence, and accumulates heuristics without +copying a numerical library or the repository's MPS methodology cards. + +## What is implemented + +```text +experiment definition + │ strict JSON, no implicit fields + ▼ +experiment-v1 + exact capability/backend binding + │ human ratifies physics and gate before compute + ▼ +external TN worker (not bundled here) + │ primary + repeat request-bound artifact chains + ▼ +evidence-v1 ──> stdlib-only gate.py + │ strict parse and identity binding + │ deterministic reconstruction of both backend bundles + │ locked reference + structural repeat metric derivation + │ raw-only diagnostics remain reported-only + ▼ + fixture contract closure or stable rejection code + │ + ▼ + append-only heuristics Library +``` + +The public capsule contains: + +- [`contracts/experiment-v1.schema.json`](contracts/experiment-v1.schema.json): + problem, physics, capability, backend binding, explicit numerics, + observables, validator assurance policies, acceptance policy, a + preregistered energy reference, and provenance. +- [`contracts/evidence-v1.schema.json`](contracts/evidence-v1.schema.json): + immutable experiment identity, matching binding, execution state, + artifact manifest, observable evidence, validator results, and result + identity. +- [`contracts/backend-result-v1.schema.json`](contracts/backend-result-v1.schema.json): + the public JSON form of the main repository's + `tn_agent.backends.models.BackendResultBundleV1`, including its semantic + canonical digest. +- [`contracts/validator-evidence-v1.schema.json`](contracts/validator-evidence-v1.schema.json): + the separately addressed gate-derived validator report. It binds the + canonical request, both semantic backend-result digests, and the + preregistered reference artifact. +- [`contracts/energy-reference-v1.schema.json`](contracts/energy-reference-v1.schema.json): + the registered reference record bound to exact physics, value, + normalization, method label, citation, and semantic identity. The gate + verifies this record but does not rerun the cited reference method. +- [`gate.py`](gate.py): a dependency-free Python CLI that rejects duplicate + keys, non-finite values, missing/unknown fields, unsupported routes, + mismatched bindings/digests, unsafe artifacts, stale normalized results, + reused primary/repeat identities, self-contradicting raw metrics, forged + validator reports, hardlinked files, and failed thresholds. +- [`library/heuristics.jsonl`](library/heuristics.jsonl): append-only, + revisioned heuristics with source, evidence, confidence, contradiction, and + supersession fields. +- [`skill/tn-agent-workflow/SKILL.md`](skill/tn-agent-workflow/SKILL.md): a + thin workflow skill that composes the repository's MPS/TeNPy skills. +- [`fixtures/`](fixtures/) and [`tests/test_gate.py`](tests/test_gate.py): + compact synthetic finite/infinite examples and negative contract tests. +- [`fixtures/regenerate.py`](fixtures/regenerate.py): deterministic + regeneration of every synthetic artifact and digest. +- [`calibration/`](calibration/): the public input manifest, five blind + generator outputs, weak controls, and deterministic report for the + registered QuantumBFS #124–#128 calibration. Sealed issue statements are + excluded. +- [`issue133-campaign/`](issue133-campaign/): five new frozen tensor-network + problems, their separately preregistered exact gates, Solver certificates, + fresh Verifier subprocess receipts, rejected negative controls, named human + acceptance decisions, a readable report, and a SHA-256 manifest. These are + distinct from the #124–#128 calibration set. + +The schemas are useful for editors and other agents. `gate.py` is the +executable semantic authority: JSON Schema alone does not compare two +documents, recompute file digests, or evaluate preregistered thresholds. + +## Current scientific scope + +Only two exact, reviewable routes are promoted: + +| Capability | Physics and algorithm | Binding | Maturity | +|---|---|---|---| +| `tenpy.finite_1d.dmrg` | finite spin-1/2 TFIM, open chain, two-site MPS DMRG | `tn-agent.tenpy.finite-tfim-dmrg.v1` → `tenpy.v1` → `tenpy` | stable | +| `tenpy.infinite_1d.vumps` | infinite spin-1/2 XXZ, two-site uniform MPS VUMPS, total Sz=0 | `tn-agent.tenpy.infinite-xxz-vumps.v1` → `tenpy.v1` → `tenpy` | experimental; variance is explicitly backend-limited | + +Both routes must request `energy` and `variance`, because those observables +are dependencies of the public gate. Finite experiments do not expose +`min_sweeps` or `entropy_tolerance`; the exact translator supplies the +backend-fixed values `0` and `null`. Infinite experiments preregister both and +must satisfy +`finite_entanglement_fit.max_chi == max_bond_dim == chi_schedule[-1]`. + +The infinite route uses +`H = Σ_i[Jxy(Sx_i Sx_{i+1}+Sy_i Sy_{i+1}) + Jxy·Delta Sz_i Sz_{i+1} - h Sz_i]`. +Thus `Delta` is dimensionless and the TeNPy coupling is +`Jz = Jxy·Delta`. This standalone capsule currently promotes only `Jxy=1`: +the corrected general mapping lives in the external TN-Agent worker, which is +not part of this PR's independently reviewable trust root. Non-unit `Jxy` +therefore returns `UNSUPPORTED_ROUTE` until a versioned worker implementation +or trusted execution receipt is included in the public evidence boundary. + +Installed packages, catalog entries, and fallback names do not create +capabilities. A request for quimb, YASTN, ITensor, MPSKit, PEPSKit, a different +model, boundary, sector, or algorithm returns `UNSUPPORTED_ROUTE`. + +Method knowledge stays in the upstream harness: + +- [Method MPS](../../../../skills/method-mps/SKILL.md) owns algorithm choice, + accuracy knobs, and scientific validation guidance. +- [Using TeNPy](../../../../skills/using-tenpy/SKILL.md) owns TeNPy setup and + API-specific workflow. + +This solution links those sources and records their SHA-256 identities in +fixtures and Library records; it does not duplicate their method content. + +## Run the public gate + +From the repository root: + +```bash +TN_PUBLIC_ROOT=tracks/agent-kb/solutions/WangTheoPhys + +python3 "$TN_PUBLIC_ROOT/gate.py" validate \ + "$TN_PUBLIC_ROOT/fixtures/valid-finite/experiment.json" + +python3 "$TN_PUBLIC_ROOT/gate.py" evaluate \ + "$TN_PUBLIC_ROOT/fixtures/valid-infinite/experiment.json" \ + "$TN_PUBLIC_ROOT/fixtures/valid-infinite/evidence.json" \ + --artifact-root "$TN_PUBLIC_ROOT/fixtures/valid-infinite/artifacts" + +python3 "$TN_PUBLIC_ROOT/gate.py" validate-library \ + "$TN_PUBLIC_ROOT/library/heuristics.jsonl" +``` + +Every operational command writes exactly one compact JSON object to stdout. +Standard argparse `--help` is the sole human-readable exception. Exit status +`0` means the requested validation/evaluation passed, `2` means an input or +contract was invalid, and `3` means well-shaped evidence did not satisfy the +preregistered experiment. Every exit-3 JSON result includes +`"accepted": false`. + +Run the direct test suite: + +```bash +python3 -m unittest discover \ + -s tracks/agent-kb/solutions/WangTheoPhys/tests -v +``` + +The `test_fixture` documents are synthetic contract tests. Their +`ACCEPTANCE_PASSED` verdict demonstrates contract-level closure through +eleven required artifacts: one +request; primary and repeat raw/normalized result pairs; a preregistered +energy-reference record; validator evidence; and primary/repeat stdout and +stderr. Both request fixtures are parsed by the main repository's +`parse_tenpy_request_json`; all four normalized fixtures are parsed by +`BackendResultBundleV1.model_validate_json` and its Python-mode +`model_validate` path, and both requests pass the main worker's route-specific +request validation. The synthetic primary and repeat records are generated by +the same fixture script, so they test structural separation only. They are not +independent numerical runs, a new physics result, or challenge-tier evidence. + +A `candidate` remains a valid frozen experiment definition, but this capsule +returns `SCIENTIFIC_EVIDENCE_UNATTESTED` for candidate evaluation because it +does not bundle a trusted runner receipt or independently checkable state +certificate. Rebuilding all candidate artifacts and digests cannot turn +worker assertions into scientific acceptance. + +## Identity and acceptance + +All identities use lowercase `sha256:<64 hex>`. + +- `problem.status=test_fixture` is required for fixture-level + `ACCEPTANCE_PASSED`. A validated `candidate` is always scientifically + rejected as unattested by this contract version. +- `experiment_digest` is SHA-256 over canonical UTF-8 JSON for the complete + experiment (`sort_keys=true`, compact separators, no NaN/Infinity). +- The outer evidence `result_digest`, both main-model backend + `result_digest` values, the energy-reference `result_digest`, and + validator-evidence `result_digest` each use the same canonical algorithm + over their respective object with only that object's `result_digest` + omitted. They are semantic identities, not raw-file identities. +- Artifact digests cover the raw file bytes. Evaluation pins one non-symlink + artifact-root directory descriptor, opens every relative component without + following symlinks, rejects files whose hard-link count is not exactly one, + and recomputes bounded regular files' size and digest. +- Every observable points to the primary normalized backend-result artifact. + Every validator result points to the separate validator-evidence artifact. +- The gate reconstructs the exact main-repository `BackendResultBundleV1` + twice, from one canonical request and two distinct raw/execution/stream + chains. Both submitted normalized bundles must be canonical-identical to + those reconstructions. Execution handles and raw-result identities must + differ. This proves artifact-level separation; it cannot prove that two + external processes were scheduled independently. +- The experiment locks the energy value, units, normalization, and byte + digest of a separate reference record. The record also binds the exact + physics digest. `benchmark_delta` is derived only from primary energy versus + that locked reference; a worker-supplied benchmark value has no authority. +- `reproduction_delta` is derived only from primary versus repeat raw energy. + It cannot be supplied in the primary raw result. It is a structural, + `reported_only` diagnostic with no operator or threshold and does not + contribute to `ACCEPTANCE_PASSED`. +- Energy drift is derived from the last two primary convergence energies. A + reported drift, when present, must agree, but it does not create an + independent physics certificate. +- Variance, canonical residual, and symmetry residual cannot be recomputed + without a state or certificate. They therefore have `reported_only` status + (or `backend_limited` for infinite-route variance), carry no acceptance + threshold, and are excluded from `all_required`. +- Only parse consistency, convergence, reference comparison, and artifact + completeness are `required_pass`. They must report `pass` and satisfy the + exact preregistered threshold. + +Reproduction may be promoted in a future contract only when the experiment +preregisters a distinct attempt nonce and runner identity and a trusted +scheduler signature/MAC or external registry receipt binds those fields to +both the experiment and request digests. Byte differences, warning text, +self-authored handles, or a second locally generated bundle do not satisfy +that requirement. + +There is no fallback parser, route, backend, validator, artifact, or success +classification. The public outcome vocabulary is documented in +[`contracts/reason-codes.md`](contracts/reason-codes.md). + +## Accumulating Library + +[`library/README.md`](library/README.md) defines the append protocol. Records +are never edited or deleted. A correction is a new consecutive revision that +names the immediately prior revision in `supersedes`; a disagreement names +only earlier records in `contradicts`. The gate derives the effective latest +record without erasing history. + +Every source and evidence entry carries an addressable, kind-checked `uri` +plus SHA-256. Repository skills and method/workflow cards must be exact +`skills//SKILL.md` paths relative to the repository root; +contract audits must be regular files below this team's `docs/` or `tests/` +directory. The gate opens and hashes those paths through confined directory +descriptors. +These checks establish local content identity, not historical immutability; +published Library state must also freeze an external Git commit/tip or +equivalent registry identity. + +The seed entries are grounded workflow heuristics, not benchmark results. +Future solver attempts—success or failure—should append evidence-bearing +records so the growth curve is auditable. + +## Status against issue #133 + +| Issue #133 requirement | Status in this PR | +|---|---| +| Versioned candidate problem with executable gate | **Candidate preregistration is implemented; candidate scientific acceptance is intentionally rejected without attestation** | +| Pre-registered, machine-checkable acceptance | **Implemented only for synthetic fixture contract closure, not a fresh candidate solve** | +| Provenance and reproducible evidence identities | **Implemented at contract/artifact level; trusted execution identity is not implemented** | +| Accumulating heuristics Library | **Schema, append protocol, validator, and seed records implemented** | +| Literature-mining problem generator | **Implemented in the AGPL-3.0-only TN-Agent control plane; this capsule publishes its registered calibration evidence, not the runtime** | +| Calibration against challenges #124–#128 | **Passed the registered self-report/grouping metrics: gap 1.0, executable gate 1.0, strong/weak separation 0.7877, and 4 of 5 literature groups recovered for 0.8** | +| Five new human-accepted challenge problems (Tier 1) | **5 / 5 submission evidence published in `issue133-campaign/`; accepted by `human.junkaiwang` as `human expert supervision`; upstream catalog determination remains with QuantumBFS maintainers** | +| Five fresh solved gates (Tier 2) | **5 / 5 exact gates pass in fresh Verifier subprocesses; all 5 registered negative controls are rejected** | +| Refereed publication (Tier 3) | **Not claimed** | + +The supervised live campaign reports complete Tier 1 and Tier 2 submission +evidence. Its five items are new and do not count the historical calibration +set. Every item binds the frozen challenge, gate, Solver certificate, +Verifier source, positive process result, rejected negative control, and named +human acceptance decision. QuantumBFS maintainers retain final authority over +upstream catalog inclusion and tier determination. No Tier 3 or refereed +publication result is claimed. + +The calibration's meaningful-gap, executable-gate, attack-pass, novelty, and +publishability inputs are self-reported candidate fields; the deterministic +evaluator checks their schema and arithmetic rather than independently +adjudicating the summaries. The public artifacts reproduce all file/semantic +digests and scores, but they do not independently prove the operator-recorded +blind chronology or generator isolation because no external signed or +timestamped blind-run receipt is included. + +Replay the live campaign and its direct tests from the repository root: + +```bash +python3 tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/run_campaign.py +python3 -m unittest discover \ + -s tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/tests -v +``` + +The current supervised counters are `5/5` accepted new problems and `5/5` +exact solved gates. The next scientific milestone is independent QuantumBFS +catalog review followed by manuscript development and refereed review. + +## Scope and licensing + +No third-party source code, model weights, papers, or numerical artifacts are +vendored here; the compact numerical files are synthetic contract fixtures. +This directory follows the upstream `quantum.harness` repository terms and +does not add a standalone or conflicting license. See +[`THIRD_PARTY_NOTICES.md`](THIRD_PARTY_NOTICES.md) for referenced external +projects and repository documents. diff --git a/tracks/agent-kb/solutions/WangTheoPhys/THIRD_PARTY_NOTICES.md b/tracks/agent-kb/solutions/WangTheoPhys/THIRD_PARTY_NOTICES.md new file mode 100644 index 000000000..3c99a1ac5 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/THIRD_PARTY_NOTICES.md @@ -0,0 +1,21 @@ +# Third-party and source notices + +This solution vendors no third-party source code, model weights, or papers. +Its small energy-reference records are original synthetic-fixture inputs, +not copied third-party numerical result files. + +It references, but does not copy: + +- `skills/method-mps/SKILL.md` and `skills/using-tenpy/SKILL.md` from this + Quantum Many-Body Physics Harness repository. +- [TeNPy](https://github.com/tenpy/tenpy), an external tensor-network package + distributed under Apache-2.0. The package is not bundled or imported by the + public gate. +- The separately maintained TN-Agent source checkout used by the optional + compatibility tests, distributed under AGPL-3.0. No TN-Agent source is + copied into this capsule. + +The contract and gate files in this team directory are original contribution +files submitted under the upstream `quantum.harness` repository terms. This +capsule does not add a standalone license or terms that conflict with either +the upstream repository or referenced projects. diff --git a/tracks/agent-kb/solutions/WangTheoPhys/calibration/.gitignore b/tracks/agent-kb/solutions/WangTheoPhys/calibration/.gitignore new file mode 100644 index 000000000..82c6a5f0b --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/calibration/.gitignore @@ -0,0 +1,2 @@ +sealed/* +!sealed/README.md diff --git a/tracks/agent-kb/solutions/WangTheoPhys/calibration/README.md b/tracks/agent-kb/solutions/WangTheoPhys/calibration/README.md new file mode 100644 index 000000000..92c1e1e39 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/calibration/README.md @@ -0,0 +1,59 @@ +# Registered blind calibration + +This directory publishes the non-secret evidence from the 2026-07-29 +QuantumBFS issue #124–#128 calibration of the TN problem generator. + +The generator received only the source-snapshot date, ten literature +identifiers, target count, and four scoring dimensions. The operator record +states that it did not receive issue statements, links, target identifiers, +statement digests, literature grouping, repository access, or earlier +discussion, and that generation completed before the five statements were +supplied to a separate evaluator. These public files bind content and +arithmetic; they do not independently prove that chronology or generator +isolation because no external signed or timestamped blind-run receipt exists. + +Files: + +- `manifest.json`: registered sources, dimensions, statement digests, and + thresholds; +- `blind-candidates.json`: the five public generator outputs; +- `weak-controls.json`: the three public weak controls used by separation; +- `report.json`: deterministic scores and hidden matching results. + +The local `.gitignore` rejects `sealed/*` except a possible explanatory +`sealed/README.md`; the test suite also allowlists every committed calibration +file so nested statement payloads fail review. + +The candidate file has SHA-256 +`101bf35d22607b08ad0b160b893392e021bf7f8ac9d54aaefce91007f3b37be8`. +The semantic report digest is +`sha256:ee4acc0b035c0678d05997fa6d6a5991a99e708d42e30d4c5caea90fb519e85f`. + +The report passed all registered campaign thresholds. Its `0.8` matching +score means that 4 of 5 literature groups met the per-target `0.75` threshold; +it is not a 5-of-5 semantic comparison with the sealed issue prose. + +- meaningful-gap recovery: `1.0`; +- executable/non-gameable gate: `1.0`; +- strong/weak separation: `0.7876666666666667`; +- blind literature-group recovery: `0.8` (`4 of 5`). + +The first three values are arithmetic over self-reported candidate fields. +Meaningful-gap recovery reads `meaningful_gap`; executable/non-gameable gate +reads `gate_executable && gate_attack_passed`; separation compares the +candidate-supplied novelty/publishability scores with `weak-controls.json`. +The evaluator checks schemas, digests, allowed literature, and arithmetic; it +does not independently verify those scientific self-reports. + +An unregistered preflight found that ASCII token matching was +language-dependent. Before the registered run, matching was changed to exact +set-Jaccard over each candidate's required source-literature identifiers and +the per-target threshold was raised from `0.16` to `0.75`. The campaign-level +`0.8` threshold did not change. + +Passing this calibration shows that the five outputs satisfied the registered +self-report and literature-grouping metrics. It does not count as a +human-accepted new challenge, a fresh numerical solve, or a publication. The +sealed issue statements are intentionally absent; their digests and the +post-run target-to-literature grouping remain public so the digest and score +chain can be replayed. diff --git a/tracks/agent-kb/solutions/WangTheoPhys/calibration/blind-candidates.json b/tracks/agent-kb/solutions/WangTheoPhys/calibration/blind-candidates.json new file mode 100644 index 000000000..8c762c51a --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/calibration/blind-candidates.json @@ -0,0 +1 @@ +[{"candidate_id":"calibration.candidate.kagome_certified_bracket","literature_ids":["doi:10.1126/science.1201080","arxiv:2212.03014"],"summary":"科学缺口:doi:10.1126/science.1201080 以圆柱 DMRG 为 kagome 自旋液体提供强数值证据,但变分上界、有限尺寸效应与热力学极限自旋隙仍未被同一可审计证书闭合。TN 中心方法:在 SU(2) 对称 MPS/PEPS 上同时构造基态与最低三重态的严格变分上界,并以 arxiv:2212.03014 的重整化约化密度矩阵 SDP 给出匹配下界;对边界扭转和周长序列用区间收缩传播截断误差。可证伪输出:每个实例提交 E0、E1 的上下区间、Δ=E1−E0 的区间、能量方差、相关长度、每步 Schmidt 尾重与可重放证书;任一区间不包含独立 ED 值、下界高于上界或声称有隙而 Δ 下端不为正即失败。预注册固定阈值策略:公开小簇上先用独立 ED 冻结允许能量区间宽度 εE 与截断尾误差系数;相判据冻结为原论文报告的能量/隙区间经 Weyl 扰动界扩张后的区间,测试后不得按所得解移动阈值。fresh input:代码与阈值封存后,由公开哈希随机种子生成未见过的 kagome 圆柱周长、边界扭转和满足预定范数上限的弱键扰动。攻击/弱解区分:只给低变分能量、只做单个周长、把不同拓扑扇区误作激发态或事后挑选扭转均不能同时通过双边界、方差、扇区正交和全种子覆盖检查;强解必须闭合非零隙区间并显示随周长稳定。允许的受审工具:小簇 ED、约化密度矩阵 SDP、区间算术、SU(2) 符号代数、MPS/PEPS 收缩及其显式时间/峰值内存/键维复杂度核算。","meaningful_gap":true,"gate_executable":true,"gate_attack_passed":true,"novelty_score":0.9,"publishability_score":0.94},{"candidate_id":"calibration.candidate.aklt_gap_certificate","literature_ids":["doi:10.1103/PhysRevLett.59.799","doi:10.1103/PhysRevLett.124.177204"],"summary":"科学缺口:doi:10.1103/PhysRevLett.59.799 给出一维 AKLT 严格有隙范例,doi:10.1103/PhysRevLett.124.177204 对二维六角格依靠 36 自旋高精度 DMRG 与有限尺寸判据得到热力学极限 Δ>0.006;缺少可移植、机器可核验并能覆盖加权邻域的 TN 证书。TN 中心方法:从 AKLT PEPS 的虚拟自旋构造精确局域投影器,以对称 MPS 求加权 36 点 cut-out Hamiltonian 的最低非零本征值,并用残差、Temple 型界和区间算术把数值结果升级为严格区间,再代入论文有限尺寸不等式。可证伪输出:对每组权重输出局域隙区间 γF、解析阈值 τ、诱导的无限体积隙下界、投影器恒等式残差和证书哈希;只要 inf γF≤τ、全局下界≤0、或 a=1.4 时不能复现 γF≈0.14599 超过 τ=0.138 并导出至少 0.006 的基准,即否证。预注册固定阈值策略:τ(a) 完全冻结为论文有限尺寸定理的解析表达式,数值安全裕量固定为区间下端减 τ(a),基准门槛固定为论文的 0.138、0.14599 和 0.006,不引用未来求解结果。fresh input:实现封存后用承诺种子抽取预先限定权重盒中的有理权重及等价但不同张量规范、站点编号和切割嵌入。攻击/弱解区分:硬编码 a=1.4、仅报告浮点 DMRG gap、遗漏基态流形投影或利用规范/编号泄漏会在新权重、精确投影器检查和区间残差中失败;强解须对整个抽样权重集给出逐实例严格证书。允许的受审工具:小子块 ED、SDP 辅助下界、区间本征值算法、SU(2) 符号运算、PEPS/MPS 精确及截断收缩、收缩复杂度审计。","meaningful_gap":true,"gate_executable":true,"gate_attack_passed":true,"novelty_score":0.86,"publishability_score":0.93},{"candidate_id":"calibration.candidate.neural_tn_separation","literature_ids":["arxiv:2211.05504","arxiv:2302.01941","arxiv:2212.03014"],"summary":"科学缺口:arxiv:2211.05504 与 arxiv:2302.01941 展示 ViT/深层 NQS 在受挫自旋模型中的高精度能量和潜在无隙自旋液体证据,但低能量本身不能排除错误的相、扇区或长程关联;需要把神经态主张转化为可独立验证的 TN 证书。TN 中心方法:对提交的 NQS 做自适应幅度查询并用 TT-cross/MPS 蒸馏,在固定最大键维下区间收缩 NQS–MPS 重叠、Hamiltonian 矩、结构因子和长距关联;再用 arxiv:2212.03014 的 RG-SDP 下界与对称 MPS 激发态上界形成能量及自旋隙双边括号。可证伪输出:每个 fresh J1-J2 实例输出保真度下界、E0/Δ 区间、能量方差、两个预注册动量的结构因子和查询转录哈希;在可 ED 尺寸上任何区间漏掉真值,或大尺寸相标签与预注册有限尺寸标度检验矛盾,均失败。预注册固定阈值策略:在与测试种子独立的小格点 ED 校准集上冻结保真度下界、每站点能量误差和结构因子误差容限;资源上限冻结为论文展示的最多 10^6 参数和 64 层,强弱分界冻结为同资源浅 ViT/MPS 基线的盲测误差置信上界,而非未来最佳解。fresh input:封存后抽取未见过的 J2/J1 有理点、方形或三角形格点、边界扭转和对称扇区,并在提交完成后才公布幅度查询构型。攻击/弱解区分:记忆已发表能量、只优化能量、不实现复振幅符号或用参数量堆叠均无法通过事后幅度查询、重叠下界、结构因子和 gap bracket;强解须在相同资源预算下同时胜过独立基线并给出相一致证据。允许的受审工具:受限尺寸 ED、RG-SDP、区间统计与区间收缩、晶格/群对称符号检查、MPS/TT/PEPS 收缩及查询数、键维、FLOP 和内存审计。","meaningful_gap":true,"gate_executable":true,"gate_attack_passed":true,"novelty_score":0.92,"publishability_score":0.9},{"candidate_id":"calibration.candidate.trotter_tn_witness","literature_ids":["arxiv:1912.08854","doi:10.1103/PhysRevLett.128.210501"],"summary":"科学缺口:arxiv:1912.08854 给出利用交换结构的通用 Trotter 误差理论,doi:10.1103/PhysRevLett.128.210501 解释一阶公式的异常低误差,但解析上界对具体多体局域可观测量何时紧、何时被干涉抵消仍缺少可重复的实例级证据。TN 中心方法:把精确时间演化和一至四阶 product formula 编译为共享因果锥 TN,以多时间点批量收缩计算局域可观测量误差;同时以符号交换子范数生成论文型解析上界,并用区间算术认证真误差和界的比值。可证伪输出:逐实例输出误差区间、解析界、界/误差比、步数标度指数、因果锥宽度和完整收缩路径;若任一真误差上端超过理论界、Heisenberg 偶奇排序基准的复杂度估计不能保持论文所述至多约 5 倍紧度、或拟合指数的区间不含预注册阶数,即失败。预注册固定阈值策略:界值直接由两篇论文的交换子公式和输入 Hamiltonian 解析计算;5 倍基准、Trotter 阶数及舍入包络在见到测试实例前冻结,零误差点不计分且时间窗口由独立小系统 ED 基线固定。fresh input:实现封存后用承诺种子生成未见过的局域 Pauli 系数、项排序、系统长度、时间点和局域观测量,另含预注册的交换与反交换压力实例。攻击/弱解区分:选择极小时间、只拟合渐近斜率、返回宽松范数界或仅跑小系统会被固定时间窗、界紧度、全尺寸因果锥与逐点区间检查拒绝;强解必须同时保证界有效、常数紧且解释排序相关抵消。允许的受审工具:小尺寸 ED、区间矩阵指数/区间算术、Pauli 与嵌套交换子符号化简、MPO/时空 TN 收缩、树宽/FLOP/峰值内存复杂度证书;不允许用未审计黑盒量子模拟器替代证书。","meaningful_gap":true,"gate_executable":true,"gate_attack_passed":true,"novelty_score":0.88,"publishability_score":0.91},{"candidate_id":"calibration.candidate.xeb_auditable_contraction","literature_ids":["arxiv:2108.05665","arxiv:2002.01935"],"summary":"科学缺口:arxiv:2108.05665 将 53 比特 Sycamore 数据的精确 XEB 推到 16 cycles,arxiv:2002.01935 表明超优化路径可带来数量级乃至约 10^4 倍的成本改善;然而现有结果通常只给最终 XEB 或启发式成本,不能同时证明振幅正确、复用合法和路径接近实例下界。TN 中心方法:把随机电路构造成带输出切片的 TN,联合使用 multi-tensor memoization 与超优化收缩路径;对每个缓存子网提交边界张量哈希,以树分解/切割 SDP 下界约束 contraction width,并对复振幅作有向舍入区间收缩。可证伪输出:输出随机位串振幅的复区间、线性 XEB 区间、缓存 DAG、路径峰值宽度、FLOP/内存和下界比值;深度≤10 的 fresh 子集必须与 ED 区间相交,16-cycle 基准必须精确完成,任一随机开切割的两侧重收缩不一致或路径成本超过冻结阈值即失败。预注册固定阈值策略:数值容限由独立 Haar/Clifford 可精确基线的有向舍入误差确定;深度门槛固定为论文的 10 与 16 cycles,速度门槛取论文公开 multi-tensor 基线和超优化基线的较弱者,路径近优阈值取提交成本除解析切割/树宽下界的固定因子 2,不按未来最优路径调整。fresh input:代码封存后由公开承诺种子生成门参数、布局缺边、位串批次和缓存压力模式,位串在路径承诺后揭示以阻止只算高概率样本。攻击/弱解区分:伪造 XEB、近似振幅、挑选容易位串、重复计算冒充缓存或只优化 FLOP 会分别被随机振幅开查、区间重收缩、后揭示位串、缓存 DAG 和峰值内存门槛捕获;强解须在同一 fresh 批次兼顾精确性、复用和近下界复杂度。允许的受审工具:小深度 ED、切割/树宽 SDP 或整数下界、区间复数算术、Clifford 符号基线、任意精确或切片 TN 收缩及可验证 FLOP/峰值内存核算。","meaningful_gap":true,"gate_executable":true,"gate_attack_passed":true,"novelty_score":0.84,"publishability_score":0.88}] diff --git a/tracks/agent-kb/solutions/WangTheoPhys/calibration/manifest.json b/tracks/agent-kb/solutions/WangTheoPhys/calibration/manifest.json new file mode 100644 index 000000000..5a31247ce --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/calibration/manifest.json @@ -0,0 +1,67 @@ +{ + "allowed_source_snapshot_date": "2026-07-29", + "calibration_id": "calibration.quantumbfs-124-128", + "digest": "sha256:c6288ad94f77779acac9de5cf9488252ff5e4461180f94d479a7f7d98193f4ae", + "dimensions": [ + "meaningful-gap-recovery", + "executable-nongameable-gate", + "strong-weak-separation", + "hidden-target-match" + ], + "schema_version": "tn-agent.factory-calibration-manifest.v1", + "targets": [ + { + "issue_url": "https://github.com/QuantumBFS/quantum.harness/issues/124", + "literature_ids": [ + "doi:10.1126/science.1201080", + "arxiv:2212.03014" + ], + "statement_digest": "sha256:94d4fa4fe643a74c6fe1f73149bf7fd604e0100a8fe7b555e5dcb2ccf65bd323", + "target_id": "quantumbfs.issue-124" + }, + { + "issue_url": "https://github.com/QuantumBFS/quantum.harness/issues/125", + "literature_ids": [ + "arxiv:2302.01941", + "arxiv:2211.05504" + ], + "statement_digest": "sha256:7886d909291de6820e5fabe974a6cb2af478a080f18e1695855ca57ca2ed4abc", + "target_id": "quantumbfs.issue-125" + }, + { + "issue_url": "https://github.com/QuantumBFS/quantum.harness/issues/126", + "literature_ids": [ + "doi:10.1103/PhysRevLett.59.799", + "doi:10.1103/PhysRevLett.124.177204" + ], + "statement_digest": "sha256:1605f6d8e5daa453cc290532b2be74176bbcab81f31c133fea7f28ccfe126841", + "target_id": "quantumbfs.issue-126" + }, + { + "issue_url": "https://github.com/QuantumBFS/quantum.harness/issues/127", + "literature_ids": [ + "arxiv:2002.01935", + "arxiv:2108.05665" + ], + "statement_digest": "sha256:19e785a54729eded1fcd7bcf6f1fce7f28b1f714c1c9fe69d52adf96704d8890", + "target_id": "quantumbfs.issue-127" + }, + { + "issue_url": "https://github.com/QuantumBFS/quantum.harness/issues/128", + "literature_ids": [ + "arxiv:1912.08854", + "doi:10.1103/PhysRevLett.128.210501" + ], + "statement_digest": "sha256:fe900176456fce9abcd4448d6d29fc47e3a42bdf353338d04a131b6bd62acdee", + "target_id": "quantumbfs.issue-128" + } + ], + "thresholds": { + "executable_gate": 0.8, + "hidden_target_match": 0.8, + "meaningful_gap_recovery": 0.8, + "minimum_match_similarity": 0.75, + "strong_weak_separation": 0.45 + }, + "weak_controls_digest": "sha256:67be7c53c4dbcd03c2c1d8f6306c631ecc65e2250f9ebde6670ecd567da05d35" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/calibration/report.json b/tracks/agent-kb/solutions/WangTheoPhys/calibration/report.json new file mode 100644 index 000000000..15d1f5d9b --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/calibration/report.json @@ -0,0 +1,60 @@ +{ + "calibration_id": "calibration.quantumbfs-124-128", + "candidate_digests": [ + "sha256:4fadcd94dbc40a54d07c4b93ec4ef3ac419727cd9b7c6066877debeadeb13738", + "sha256:7dfe5ea4bce9f6be2de5d1f81b24da03b3bc2ccdd7afd0e736ad9033195194af", + "sha256:06d76126a33796248ab9de37542ec3a77d462561898a6df21e756e3a049c89ae", + "sha256:0868427c8d4ccefbf949b9cbd6502cc7bfcb55e0e4a82a94f969f5e3e239aae5", + "sha256:95b887d585665a7ab54960b8738d0e922a8afb7655995f751a2fe60124d62c43" + ], + "digest": "sha256:ee4acc0b035c0678d05997fa6d6a5991a99e708d42e30d4c5caea90fb519e85f", + "failure_codes": [], + "manifest_digest": "sha256:c6288ad94f77779acac9de5cf9488252ff5e4461180f94d479a7f7d98193f4ae", + "matches": [ + { + "candidate_id": "calibration.candidate.kagome_certified_bracket", + "matched": true, + "similarity": 1.0, + "target_id": "quantumbfs.issue-124" + }, + { + "candidate_id": null, + "matched": false, + "similarity": 0.6666666666666666, + "target_id": "quantumbfs.issue-125" + }, + { + "candidate_id": "calibration.candidate.aklt_gap_certificate", + "matched": true, + "similarity": 1.0, + "target_id": "quantumbfs.issue-126" + }, + { + "candidate_id": "calibration.candidate.xeb_auditable_contraction", + "matched": true, + "similarity": 1.0, + "target_id": "quantumbfs.issue-127" + }, + { + "candidate_id": "calibration.candidate.trotter_tn_witness", + "matched": true, + "similarity": 1.0, + "target_id": "quantumbfs.issue-128" + } + ], + "passed": true, + "schema_version": "tn-agent.factory-calibration-report.v1", + "scores": { + "executable_gate": 1.0, + "hidden_target_match": 0.8, + "meaningful_gap_recovery": 1.0, + "strong_weak_separation": 0.7876666666666667 + }, + "sealed_target_digests": [ + "sha256:94d4fa4fe643a74c6fe1f73149bf7fd604e0100a8fe7b555e5dcb2ccf65bd323", + "sha256:7886d909291de6820e5fabe974a6cb2af478a080f18e1695855ca57ca2ed4abc", + "sha256:1605f6d8e5daa453cc290532b2be74176bbcab81f31c133fea7f28ccfe126841", + "sha256:19e785a54729eded1fcd7bcf6f1fce7f28b1f714c1c9fe69d52adf96704d8890", + "sha256:fe900176456fce9abcd4448d6d29fc47e3a42bdf353338d04a131b6bd62acdee" + ] +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/calibration/weak-controls.json b/tracks/agent-kb/solutions/WangTheoPhys/calibration/weak-controls.json new file mode 100644 index 000000000..9bba63140 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/calibration/weak-controls.json @@ -0,0 +1,17 @@ +[ + { + "control_id": "control.repeat-known-result", + "novelty_score": 0.05, + "publishability_score": 0.05 + }, + { + "control_id": "control.no-executable-gate", + "novelty_score": 0.2, + "publishability_score": 0.1 + }, + { + "control_id": "control.non-tn-central", + "novelty_score": 0.1, + "publishability_score": 0.15 + } +] diff --git a/tracks/agent-kb/solutions/WangTheoPhys/contracts/backend-result-v1.schema.json b/tracks/agent-kb/solutions/WangTheoPhys/contracts/backend-result-v1.schema.json new file mode 100644 index 000000000..b5798f97b --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/contracts/backend-result-v1.schema.json @@ -0,0 +1,463 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "https://github.com/QuantumBFS/quantum.harness/blob/main/tracks/agent-kb/solutions/WangTheoPhys/contracts/backend-result-v1.schema.json", + "title": "tn-agent BackendResultBundleV1", + "description": "Public JSON form of tn_agent.backends.models.BackendResultBundleV1. The semantic result_digest is the canonical SHA-256 of every field except result_digest.", + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "request_digest", + "binding", + "backend", + "environment", + "execution", + "observables", + "convergence", + "diagnostics", + "provenance", + "artifacts", + "warnings", + "known_limitations", + "backend_limited_fields", + "result_digest" + ], + "properties": { + "schema_version": { + "const": "tn-agent.backend-result.v1" + }, + "request_digest": { + "$ref": "#/$defs/digest" + }, + "binding": { + "$ref": "#/$defs/binding" + }, + "backend": { + "$ref": "#/$defs/backend" + }, + "environment": { + "$ref": "#/$defs/environment" + }, + "execution": { + "$ref": "#/$defs/execution" + }, + "observables": { + "type": "object", + "minProperties": 1, + "maxProperties": 4096, + "propertyNames": { + "$ref": "#/$defs/identifier" + }, + "additionalProperties": { + "$ref": "#/$defs/observable" + } + }, + "convergence": { + "type": "array", + "maxItems": 4096, + "items": { + "$ref": "#/$defs/convergence_point" + } + }, + "diagnostics": { + "type": "array", + "maxItems": 4096, + "items": { + "$ref": "#/$defs/diagnostic" + } + }, + "provenance": { + "$ref": "#/$defs/provenance" + }, + "artifacts": { + "type": "array", + "maxItems": 4096, + "items": { + "$ref": "#/$defs/result_artifact" + } + }, + "warnings": { + "$ref": "#/$defs/string_array" + }, + "known_limitations": { + "$ref": "#/$defs/string_array" + }, + "backend_limited_fields": { + "type": "array", + "maxItems": 4096, + "items": { + "$ref": "#/$defs/identifier" + } + }, + "result_digest": { + "$ref": "#/$defs/digest" + } + }, + "$defs": { + "identifier": { + "type": "string", + "pattern": "^[a-z0-9][a-z0-9_.-]{0,127}$" + }, + "digest": { + "type": "string", + "pattern": "^sha256:[0-9a-f]{64}$" + }, + "nonempty": { + "type": "string", + "minLength": 1, + "maxLength": 4096 + }, + "nullable_nonempty": { + "oneOf": [ + { + "$ref": "#/$defs/nonempty" + }, + { + "type": "null" + } + ] + }, + "string_array": { + "type": "array", + "maxItems": 4096, + "items": { + "$ref": "#/$defs/nonempty" + } + }, + "json_value": { + "oneOf": [ + { + "type": "null" + }, + { + "type": "boolean" + }, + { + "type": "string" + }, + { + "type": "number" + }, + { + "type": "array", + "items": { + "$ref": "#/$defs/json_value" + } + }, + { + "type": "object", + "additionalProperties": { + "$ref": "#/$defs/json_value" + } + } + ] + }, + "binding": { + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "capability_id", + "adapter_id", + "backend_id", + "request_schema", + "result_schema" + ], + "properties": { + "schema_version": { + "const": "tn-agent.backend-binding.v1" + }, + "capability_id": { + "$ref": "#/$defs/identifier" + }, + "adapter_id": { + "$ref": "#/$defs/identifier" + }, + "backend_id": { + "$ref": "#/$defs/identifier" + }, + "request_schema": { + "type": "string", + "pattern": "^tn-agent\\.[a-z0-9.-]+\\.v1$" + }, + "result_schema": { + "const": "tn-agent.backend-result.v1" + } + } + }, + "backend": { + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "backend_id", + "backend_version", + "adapter_id", + "adapter_version" + ], + "properties": { + "schema_version": { + "const": "tn-agent.backend-identity.v1" + }, + "backend_id": { + "$ref": "#/$defs/identifier" + }, + "backend_version": { + "$ref": "#/$defs/nonempty" + }, + "adapter_id": { + "$ref": "#/$defs/identifier" + }, + "adapter_version": { + "$ref": "#/$defs/nonempty" + } + } + }, + "environment": { + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "environment_digest", + "runtime_id", + "runtime_version", + "dependencies" + ], + "properties": { + "schema_version": { + "const": "tn-agent.environment-identity.v1" + }, + "environment_digest": { + "$ref": "#/$defs/digest" + }, + "runtime_id": { + "$ref": "#/$defs/identifier" + }, + "runtime_version": { + "$ref": "#/$defs/nonempty" + }, + "dependencies": { + "type": "object", + "maxProperties": 4096, + "propertyNames": { + "$ref": "#/$defs/identifier" + }, + "additionalProperties": { + "type": "string" + } + } + } + }, + "execution": { + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "status", + "return_code", + "execution_handle", + "retryable", + "stdout_digest", + "stderr_digest", + "stdout_truncated", + "stderr_truncated" + ], + "properties": { + "schema_version": { + "const": "tn-agent.execution-evidence.v1" + }, + "status": { + "enum": [ + "succeeded", + "failed", + "timed_out", + "cancelled", + "not_executed" + ] + }, + "return_code": { + "type": [ + "integer", + "null" + ] + }, + "execution_handle": { + "$ref": "#/$defs/nullable_nonempty" + }, + "retryable": { + "type": "boolean" + }, + "stdout_digest": { + "$ref": "#/$defs/digest" + }, + "stderr_digest": { + "$ref": "#/$defs/digest" + }, + "stdout_truncated": { + "type": "boolean" + }, + "stderr_truncated": { + "type": "boolean" + } + } + }, + "observable": { + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "name", + "status", + "value", + "units", + "normalization", + "reason" + ], + "properties": { + "schema_version": { + "const": "tn-agent.observable-evidence.v1" + }, + "name": { + "$ref": "#/$defs/identifier" + }, + "status": { + "enum": [ + "measured", + "derived", + "unavailable", + "backend_limited" + ] + }, + "value": { + "$ref": "#/$defs/json_value" + }, + "units": { + "$ref": "#/$defs/nullable_nonempty" + }, + "normalization": { + "$ref": "#/$defs/nullable_nonempty" + }, + "reason": { + "$ref": "#/$defs/nullable_nonempty" + } + } + }, + "convergence_point": { + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "step", + "metrics" + ], + "properties": { + "schema_version": { + "const": "tn-agent.convergence-point.v1" + }, + "step": { + "type": "integer", + "minimum": 0 + }, + "metrics": { + "type": "object", + "maxProperties": 4096, + "propertyNames": { + "$ref": "#/$defs/identifier" + }, + "additionalProperties": { + "type": "number" + } + } + } + }, + "diagnostic": { + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "code", + "severity", + "message", + "field" + ], + "properties": { + "schema_version": { + "const": "tn-agent.diagnostic.v1" + }, + "code": { + "$ref": "#/$defs/identifier" + }, + "severity": { + "enum": [ + "info", + "warning", + "error" + ] + }, + "message": { + "$ref": "#/$defs/nonempty" + }, + "field": { + "$ref": "#/$defs/nullable_nonempty" + } + } + }, + "provenance": { + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "plan_id", + "capability_id", + "adapter_id", + "raw_result_relative" + ], + "properties": { + "schema_version": { + "const": "tn-agent.provenance-evidence.v1" + }, + "plan_id": { + "$ref": "#/$defs/digest" + }, + "capability_id": { + "$ref": "#/$defs/identifier" + }, + "adapter_id": { + "$ref": "#/$defs/identifier" + }, + "raw_result_relative": { + "type": "string" + } + } + }, + "result_artifact": { + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "relative_path", + "digest", + "media_type", + "size_bytes" + ], + "properties": { + "schema_version": { + "const": "tn-agent.result-artifact.v1" + }, + "relative_path": { + "type": "string" + }, + "digest": { + "$ref": "#/$defs/digest" + }, + "media_type": { + "$ref": "#/$defs/nonempty" + }, + "size_bytes": { + "type": "integer", + "minimum": 0 + } + } + } + } +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/contracts/energy-reference-v1.schema.json b/tracks/agent-kb/solutions/WangTheoPhys/contracts/energy-reference-v1.schema.json new file mode 100644 index 000000000..b0246a123 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/contracts/energy-reference-v1.schema.json @@ -0,0 +1,70 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "https://github.com/QuantumBFS/quantum.harness/blob/main/tracks/agent-kb/solutions/WangTheoPhys/contracts/energy-reference-v1.schema.json", + "title": "WangTheoPhys Preregistered Energy Reference v1", + "description": "A preregistered energy anchor with exact physics and semantic identities. The gate verifies the record; it does not rerun the cited reference method.", + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "reference_id", + "capability_id", + "physics_digest", + "observable", + "value", + "units", + "normalization", + "method", + "citation", + "result_digest" + ], + "properties": { + "schema_version": { + "const": "wangtheophys.tn-energy-reference.v1" + }, + "reference_id": { + "$ref": "#/$defs/identifier" + }, + "capability_id": { + "$ref": "#/$defs/identifier" + }, + "physics_digest": { + "$ref": "#/$defs/digest" + }, + "observable": { + "const": "energy" + }, + "value": { + "type": "number" + }, + "units": { + "const": "J" + }, + "normalization": { + "enum": [ + "total", + "per-site" + ] + }, + "method": { + "$ref": "#/$defs/identifier" + }, + "citation": { + "type": "string", + "minLength": 1 + }, + "result_digest": { + "$ref": "#/$defs/digest" + } + }, + "$defs": { + "identifier": { + "type": "string", + "pattern": "^[A-Za-z0-9][A-Za-z0-9_.:@-]{0,127}$" + }, + "digest": { + "type": "string", + "pattern": "^sha256:[0-9a-f]{64}$" + } + } +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/contracts/evidence-v1.schema.json b/tracks/agent-kb/solutions/WangTheoPhys/contracts/evidence-v1.schema.json new file mode 100644 index 000000000..b2986a7d8 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/contracts/evidence-v1.schema.json @@ -0,0 +1,603 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "https://github.com/QuantumBFS/quantum.harness/blob/main/tracks/agent-kb/solutions/WangTheoPhys/contracts/evidence-v1.schema.json", + "title": "WangTheoPhys Tensor-Network Evidence v1", + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "experiment_digest", + "binding", + "execution", + "repeat_execution", + "artifacts", + "observables", + "validator_results", + "provenance", + "result_digest" + ], + "properties": { + "schema_version": { + "const": "wangtheophys.tn-evidence.v1" + }, + "experiment_digest": { + "$ref": "#/$defs/digest" + }, + "binding": { + "$ref": "#/$defs/binding" + }, + "execution": { + "$ref": "#/$defs/execution" + }, + "repeat_execution": { + "$ref": "#/$defs/execution" + }, + "artifacts": { + "type": "array", + "minItems": 11, + "maxItems": 11, + "items": { + "$ref": "#/$defs/artifact" + } + }, + "observables": { + "type": "array", + "minItems": 1, + "items": { + "$ref": "#/$defs/observable" + } + }, + "validator_results": { + "type": "array", + "minItems": 1, + "items": { + "$ref": "#/$defs/validator_result" + } + }, + "provenance": { + "$ref": "#/$defs/provenance" + }, + "result_digest": { + "$ref": "#/$defs/digest" + } + }, + "allOf": [ + { + "properties": { + "artifacts": { + "contains": { + "properties": { + "role": { + "const": "backend_request" + } + }, + "required": ["role"] + }, + "minContains": 1, + "maxContains": 1 + } + } + }, + { + "properties": { + "artifacts": { + "contains": { + "properties": { + "role": { + "const": "backend_result" + } + }, + "required": ["role"] + }, + "minContains": 1, + "maxContains": 1 + } + } + }, + { + "properties": { + "artifacts": { + "contains": { + "properties": { + "role": { + "const": "backend_raw_result" + } + }, + "required": ["role"] + }, + "minContains": 1, + "maxContains": 1 + } + } + }, + { + "properties": { + "artifacts": { + "contains": { + "properties": { + "role": { + "const": "validator_evidence" + } + }, + "required": ["role"] + }, + "minContains": 1, + "maxContains": 1 + } + } + }, + { + "properties": { + "artifacts": { + "contains": { + "properties": { + "role": { + "const": "backend_stdout" + } + }, + "required": ["role"] + }, + "minContains": 1, + "maxContains": 1 + } + } + }, + { + "properties": { + "artifacts": { + "contains": { + "properties": { + "role": { + "const": "backend_stderr" + } + }, + "required": ["role"] + }, + "minContains": 1, + "maxContains": 1 + } + } + }, + { + "properties": { + "artifacts": { + "contains": { + "properties": { + "role": { + "const": "backend_repeat_raw_result" + } + }, + "required": ["role"] + }, + "minContains": 1, + "maxContains": 1 + } + } + }, + { + "properties": { + "artifacts": { + "contains": { + "properties": { + "role": { + "const": "backend_repeat_result" + } + }, + "required": ["role"] + }, + "minContains": 1, + "maxContains": 1 + } + } + }, + { + "properties": { + "artifacts": { + "contains": { + "properties": { + "role": { + "const": "energy_reference" + } + }, + "required": ["role"] + }, + "minContains": 1, + "maxContains": 1 + } + } + }, + { + "properties": { + "artifacts": { + "contains": { + "properties": { + "role": { + "const": "backend_repeat_stdout" + } + }, + "required": ["role"] + }, + "minContains": 1, + "maxContains": 1 + } + } + }, + { + "properties": { + "artifacts": { + "contains": { + "properties": { + "role": { + "const": "backend_repeat_stderr" + } + }, + "required": ["role"] + }, + "minContains": 1, + "maxContains": 1 + } + } + } + ], + "$defs": { + "identifier": { + "type": "string", + "pattern": "^[A-Za-z0-9][A-Za-z0-9_.:@-]{0,127}$" + }, + "digest": { + "type": "string", + "pattern": "^sha256:[0-9a-f]{64}$" + }, + "nonempty": { + "type": "string", + "minLength": 1 + }, + "binding": { + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "capability_id", + "adapter_id", + "backend_id", + "request_schema", + "result_schema" + ], + "properties": { + "schema_version": { + "const": "tn-agent.backend-binding.v1" + }, + "capability_id": { + "$ref": "#/$defs/identifier" + }, + "adapter_id": { + "$ref": "#/$defs/identifier" + }, + "backend_id": { + "$ref": "#/$defs/identifier" + }, + "request_schema": { + "$ref": "#/$defs/identifier" + }, + "result_schema": { + "$ref": "#/$defs/identifier" + } + } + }, + "execution": { + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "status", + "return_code", + "execution_handle", + "retryable", + "stdout_digest", + "stderr_digest", + "stdout_truncated", + "stderr_truncated" + ], + "properties": { + "schema_version": { + "const": "tn-agent.execution-evidence.v1" + }, + "status": { + "enum": [ + "succeeded", + "failed", + "timed_out", + "cancelled", + "not_executed" + ] + }, + "return_code": { + "type": [ + "integer", + "null" + ] + }, + "execution_handle": { + "type": [ + "string", + "null" + ] + }, + "retryable": { + "type": "boolean" + }, + "stdout_digest": { + "$ref": "#/$defs/digest" + }, + "stderr_digest": { + "$ref": "#/$defs/digest" + }, + "stdout_truncated": { + "type": "boolean" + }, + "stderr_truncated": { + "type": "boolean" + } + } + }, + "artifact": { + "type": "object", + "additionalProperties": false, + "oneOf": [ + { + "properties": { + "role": { + "const": "backend_request" + }, + "media_type": { + "const": "application/json" + } + }, + "required": ["role", "media_type"] + }, + { + "properties": { + "role": { + "const": "backend_raw_result" + }, + "media_type": { + "const": "application/json" + } + }, + "required": ["role", "media_type"] + }, + { + "properties": { + "role": { + "const": "backend_result" + }, + "media_type": { + "const": "application/json" + } + }, + "required": ["role", "media_type"] + }, + { + "properties": { + "role": { + "const": "validator_evidence" + }, + "media_type": { + "const": "application/json" + } + }, + "required": ["role", "media_type"] + }, + { + "properties": { + "role": { + "const": "backend_stdout" + }, + "media_type": { + "const": "text/plain" + } + }, + "required": ["role", "media_type"] + }, + { + "properties": { + "role": { + "const": "backend_stderr" + }, + "media_type": { + "const": "text/plain" + } + }, + "required": ["role", "media_type"] + }, + { + "properties": { + "role": { + "const": "backend_repeat_raw_result" + }, + "media_type": { + "const": "application/json" + } + }, + "required": ["role", "media_type"] + }, + { + "properties": { + "role": { + "const": "backend_repeat_result" + }, + "media_type": { + "const": "application/json" + } + }, + "required": ["role", "media_type"] + }, + { + "properties": { + "role": { + "const": "energy_reference" + }, + "media_type": { + "const": "application/json" + } + }, + "required": ["role", "media_type"] + }, + { + "properties": { + "role": { + "const": "backend_repeat_stdout" + }, + "media_type": { + "const": "text/plain" + } + }, + "required": ["role", "media_type"] + }, + { + "properties": { + "role": { + "const": "backend_repeat_stderr" + }, + "media_type": { + "const": "text/plain" + } + }, + "required": ["role", "media_type"] + } + ], + "required": [ + "relative_path", + "digest", + "size_bytes", + "media_type", + "role" + ], + "properties": { + "relative_path": { + "type": "string", + "pattern": "^(?!.*[\\u0000-\\u001F\\u007F])(?!/)(?![A-Za-z]:)(?!.*//)(?!.*(?:^|/)\\.{1,2}(?:/|$))(?!.*\\\\)(?!.*\\/$).+$" + }, + "digest": { + "$ref": "#/$defs/digest" + }, + "size_bytes": { + "type": "integer", + "minimum": 0, + "maximum": 1048576 + }, + "media_type": { + "$ref": "#/$defs/nonempty" + }, + "role": { + "enum": [ + "backend_request", + "backend_result", + "backend_raw_result", + "backend_repeat_result", + "backend_repeat_raw_result", + "energy_reference", + "validator_evidence", + "backend_stdout", + "backend_stderr", + "backend_repeat_stdout", + "backend_repeat_stderr" + ] + } + } + }, + "observable": { + "type": "object", + "additionalProperties": false, + "required": [ + "name", + "status", + "evidence_digest" + ], + "properties": { + "name": { + "$ref": "#/$defs/identifier" + }, + "status": { + "enum": [ + "measured", + "derived", + "backend_limited" + ] + }, + "evidence_digest": { + "$ref": "#/$defs/digest" + } + } + }, + "validator_result": { + "type": "object", + "additionalProperties": false, + "required": [ + "id", + "status", + "reason_code", + "metric_value", + "evidence_digest" + ], + "properties": { + "id": { + "$ref": "#/$defs/identifier" + }, + "status": { + "enum": [ + "pass", + "fail", + "reported_only", + "backend_limited" + ] + }, + "reason_code": { + "$ref": "#/$defs/identifier" + }, + "metric_value": { + "type": [ + "number", + "null" + ] + }, + "evidence_digest": { + "$ref": "#/$defs/digest" + } + } + }, + "provenance": { + "type": "object", + "additionalProperties": false, + "required": [ + "plan_id", + "request_digest", + "backend_result_digest", + "repeat_backend_result_digest", + "generated_by", + "generated_at" + ], + "properties": { + "plan_id": { + "$ref": "#/$defs/digest" + }, + "request_digest": { + "$ref": "#/$defs/digest" + }, + "backend_result_digest": { + "$ref": "#/$defs/digest" + }, + "repeat_backend_result_digest": { + "$ref": "#/$defs/digest" + }, + "generated_by": { + "$ref": "#/$defs/nonempty" + }, + "generated_at": { + "type": "string", + "format": "date-time" + } + } + } + } +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/contracts/experiment-v1.schema.json b/tracks/agent-kb/solutions/WangTheoPhys/contracts/experiment-v1.schema.json new file mode 100644 index 000000000..417048edf --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/contracts/experiment-v1.schema.json @@ -0,0 +1,888 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "https://github.com/QuantumBFS/quantum.harness/blob/main/tracks/agent-kb/solutions/WangTheoPhys/contracts/experiment-v1.schema.json", + "title": "WangTheoPhys Tensor-Network Experiment v1", + "description": "Fail-closed preregistration contract. The executable gate is the semantic authority for route coherence and acceptance policy.", + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "problem", + "physics", + "capability", + "backend_binding", + "numerics", + "observables", + "validators", + "acceptance", + "reference", + "provenance" + ], + "properties": { + "schema_version": { + "const": "wangtheophys.tn-experiment.v1" + }, + "problem": { + "$ref": "#/$defs/problem" + }, + "physics": { + "$ref": "#/$defs/physics" + }, + "capability": { + "$ref": "#/$defs/capability" + }, + "backend_binding": { + "$ref": "#/$defs/binding" + }, + "numerics": { + "$ref": "#/$defs/numerics" + }, + "observables": { + "type": "array", + "minItems": 1, + "uniqueItems": true, + "items": { + "$ref": "#/$defs/identifier" + }, + "allOf": [ + { + "contains": { + "const": "energy" + } + }, + { + "contains": { + "const": "variance" + } + } + ] + }, + "validators": { + "type": "array", + "minItems": 8, + "maxItems": 8, + "items": { + "$ref": "#/$defs/validator" + } + }, + "acceptance": { + "$ref": "#/$defs/acceptance" + }, + "reference": { + "$ref": "#/$defs/reference" + }, + "provenance": { + "$ref": "#/$defs/provenance" + } + }, + "$defs": { + "identifier": { + "type": "string", + "pattern": "^[A-Za-z0-9][A-Za-z0-9_.:@-]{0,127}$" + }, + "digest": { + "type": "string", + "pattern": "^sha256:[0-9a-f]{64}$" + }, + "nonempty": { + "type": "string", + "minLength": 1 + }, + "problem": { + "type": "object", + "additionalProperties": false, + "required": [ + "problem_id", + "title", + "research_question", + "novelty_claim", + "status", + "task_family" + ], + "properties": { + "problem_id": { + "$ref": "#/$defs/identifier" + }, + "title": { + "$ref": "#/$defs/nonempty" + }, + "research_question": { + "$ref": "#/$defs/nonempty" + }, + "novelty_claim": { + "$ref": "#/$defs/nonempty" + }, + "status": { + "enum": [ + "candidate", + "test_fixture" + ] + }, + "task_family": { + "enum": [ + "ground_state_1d_finite", + "ground_state_1d_infinite" + ] + } + } + }, + "model": { + "oneOf": [ + { + "$ref": "#/$defs/tfim_model" + }, + { + "$ref": "#/$defs/xxz_model" + } + ] + }, + "tfim_model": { + "type": "object", + "additionalProperties": false, + "required": [ + "representation", + "family", + "operators", + "couplings", + "onsite_terms", + "neighbor_range" + ], + "properties": { + "representation": { + "const": "operator_sum" + }, + "family": { + "const": "tfim" + }, + "operators": { + "type": "array", + "minItems": 2, + "maxItems": 2, + "uniqueItems": true, + "items": { + "enum": [ + "sigma_x_i*sigma_x_i+1", + "sigma_z_i" + ] + } + }, + "couplings": { + "type": "object", + "additionalProperties": false, + "required": [ + "J", + "g" + ], + "properties": { + "J": { + "type": "number" + }, + "g": { + "type": "number" + } + } + }, + "onsite_terms": { + "const": [] + }, + "neighbor_range": { + "const": 1 + } + } + }, + "xxz_model": { + "type": "object", + "description": "Spin-1/2 XXZ Hamiltonian H=sum_i[Jxy(Sx_i Sx_{i+1}+Sy_i Sy_{i+1})+Jxy*Delta Sz_i Sz_{i+1}-h Sz_i], so Delta is dimensionless and Jz=Jxy*Delta.", + "additionalProperties": false, + "required": [ + "representation", + "family", + "operators", + "couplings", + "onsite_terms", + "neighbor_range" + ], + "properties": { + "representation": { + "const": "operator_sum" + }, + "family": { + "const": "xxz" + }, + "operators": { + "type": "array", + "minItems": 4, + "maxItems": 4, + "uniqueItems": true, + "items": { + "enum": [ + "Sx_i*Sx_i+1", + "Sy_i*Sy_i+1", + "Sz_i*Sz_i+1", + "Sz_i" + ] + } + }, + "couplings": { + "type": "object", + "additionalProperties": false, + "required": [ + "Jxy", + "Delta", + "h" + ], + "properties": { + "Jxy": { + "type": "number", + "const": 1.0 + }, + "Delta": { + "type": "number" + }, + "h": { + "type": "number" + } + } + }, + "onsite_terms": { + "const": [] + }, + "neighbor_range": { + "const": 1 + } + } + }, + "lattice": { + "type": "object", + "additionalProperties": false, + "required": [ + "type", + "length", + "boundary", + "local_dim", + "unit_cell" + ], + "properties": { + "type": { + "const": "chain" + }, + "length": { + "type": [ + "integer", + "null" + ], + "minimum": 1 + }, + "boundary": { + "enum": [ + "open", + "infinite" + ] + }, + "local_dim": { + "const": 2 + }, + "unit_cell": { + "enum": [ + 1, + 2 + ] + } + } + }, + "symmetry": { + "type": "object", + "additionalProperties": false, + "required": [ + "U1_Sz", + "parity", + "translation", + "target_sector" + ], + "properties": { + "U1_Sz": { + "type": "boolean" + }, + "parity": { + "type": "boolean" + }, + "translation": { + "type": "boolean" + }, + "target_sector": { + "oneOf": [ + { + "type": "null" + }, + { + "$ref": "#/$defs/target_sector" + } + ] + } + } + }, + "target_sector": { + "type": "object", + "additionalProperties": false, + "required": [ + "total_Sz" + ], + "properties": { + "total_Sz": { + "type": "integer" + } + } + }, + "ansatz": { + "type": "object", + "additionalProperties": false, + "required": [ + "family", + "algorithm", + "variant", + "initial_state", + "allow_bond_growth", + "target_precision" + ], + "properties": { + "family": { + "const": "mps" + }, + "algorithm": { + "enum": [ + "dmrg", + "vumps" + ] + }, + "variant": { + "const": "two_site" + }, + "initial_state": { + "enum": [ + "all_z_plus", + "neel" + ] + }, + "allow_bond_growth": { + "const": true + }, + "target_precision": { + "type": "number", + "exclusiveMinimum": 0 + } + } + }, + "physics": { + "type": "object", + "additionalProperties": false, + "required": [ + "model", + "lattice", + "symmetry", + "ansatz" + ], + "properties": { + "model": { + "$ref": "#/$defs/model" + }, + "lattice": { + "$ref": "#/$defs/lattice" + }, + "symmetry": { + "$ref": "#/$defs/symmetry" + }, + "ansatz": { + "$ref": "#/$defs/ansatz" + } + } + }, + "capability": { + "type": "object", + "additionalProperties": false, + "required": [ + "capability_id", + "maturity", + "known_limitations" + ], + "properties": { + "capability_id": { + "enum": [ + "tenpy.finite_1d.dmrg", + "tenpy.infinite_1d.vumps" + ] + }, + "maturity": { + "enum": [ + "stable", + "experimental" + ] + }, + "known_limitations": { + "type": "array", + "minItems": 1, + "uniqueItems": true, + "items": { + "$ref": "#/$defs/nonempty" + } + } + } + }, + "binding": { + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "capability_id", + "adapter_id", + "backend_id", + "request_schema", + "result_schema" + ], + "properties": { + "schema_version": { + "const": "tn-agent.backend-binding.v1" + }, + "capability_id": { + "enum": [ + "tenpy.finite_1d.dmrg", + "tenpy.infinite_1d.vumps" + ] + }, + "adapter_id": { + "const": "tenpy.v1" + }, + "backend_id": { + "const": "tenpy" + }, + "request_schema": { + "enum": [ + "tn-agent.tenpy.finite-tfim-dmrg.v1", + "tn-agent.tenpy.infinite-xxz-vumps.v1" + ] + }, + "result_schema": { + "const": "tn-agent.backend-result.v1" + } + } + }, + "fit": { + "type": "object", + "additionalProperties": false, + "required": [ + "enabled", + "min_chi", + "max_chi" + ], + "properties": { + "enabled": { + "type": "boolean" + }, + "min_chi": { + "type": "integer", + "minimum": 1 + }, + "max_chi": { + "type": "integer", + "minimum": 1 + } + } + }, + "transfer": { + "type": "object", + "additionalProperties": false, + "required": [ + "compute", + "num_eigs" + ], + "properties": { + "compute": { + "type": "boolean" + }, + "num_eigs": { + "type": "integer", + "minimum": 1 + } + } + }, + "numerics": { + "type": "object", + "additionalProperties": false, + "required": [ + "max_bond_dim", + "cutoff", + "max_sweeps", + "mixer", + "lanczos_maxiter", + "checkpoint_every", + "finite_entanglement_fit", + "transfer_matrix", + "seed" + ], + "properties": { + "max_bond_dim": { + "type": "integer", + "minimum": 1 + }, + "cutoff": { + "type": "number", + "minimum": 0 + }, + "max_sweeps": { + "type": "integer", + "minimum": 1 + }, + "min_sweeps": { + "type": "integer", + "minimum": 1 + }, + "mixer": { + "type": "boolean" + }, + "lanczos_maxiter": { + "type": "integer", + "minimum": 1 + }, + "checkpoint_every": { + "type": "integer", + "minimum": 1 + }, + "entropy_tolerance": { + "type": "number", + "exclusiveMinimum": 0 + }, + "finite_entanglement_fit": { + "$ref": "#/$defs/fit" + }, + "transfer_matrix": { + "$ref": "#/$defs/transfer" + }, + "seed": { + "type": "integer", + "minimum": 0 + } + } + }, + "validator": { + "type": "object", + "additionalProperties": false, + "required": [ + "id", + "policy", + "metric", + "operator", + "threshold" + ], + "properties": { + "id": { + "enum": [ + "parse_consistency", + "convergence", + "variance", + "canonical_form", + "symmetry_check", + "benchmark_compare", + "reproducibility", + "artifact_completeness" + ] + }, + "policy": { + "enum": [ + "required_pass", + "reported_only", + "backend_limited" + ] + }, + "metric": { + "type": [ + "string", + "null" + ] + }, + "operator": { + "enum": [ + "max", + "min", + "equals", + null + ] + }, + "threshold": { + "type": [ + "number", + "null" + ] + } + } + }, + "acceptance": { + "type": "object", + "additionalProperties": false, + "required": [ + "mode", + "required_validator_ids", + "allowed_backend_limited_ids", + "reported_only_validator_ids", + "require_execution_success" + ], + "properties": { + "mode": { + "const": "all_required" + }, + "required_validator_ids": { + "type": "array", + "minItems": 1, + "uniqueItems": true, + "items": { + "$ref": "#/$defs/identifier" + } + }, + "allowed_backend_limited_ids": { + "type": "array", + "uniqueItems": true, + "items": { + "$ref": "#/$defs/identifier" + } + }, + "reported_only_validator_ids": { + "type": "array", + "uniqueItems": true, + "items": { + "$ref": "#/$defs/identifier" + } + }, + "require_execution_success": { + "const": true + } + } + }, + "reference": { + "type": "object", + "additionalProperties": false, + "required": [ + "observable", + "value", + "units", + "normalization", + "source" + ], + "properties": { + "observable": { + "const": "energy" + }, + "value": { + "type": "number" + }, + "units": { + "const": "J" + }, + "normalization": { + "enum": [ + "total", + "per-site" + ] + }, + "source": { + "type": "object", + "additionalProperties": false, + "required": [ + "kind", + "uri", + "sha256" + ], + "properties": { + "kind": { + "const": "registered_artifact" + }, + "uri": { + "type": "string", + "pattern": "^(?!.*[\\u0000-\\u001F\\u007F])(?!/)(?![A-Za-z]:)(?!.*//)(?!.*(?:^|/)\\.{1,2}(?:/|$))(?!.*\\\\)(?!.*\\/$).+$" + }, + "sha256": { + "$ref": "#/$defs/digest" + } + } + } + } + }, + "source": { + "type": "object", + "additionalProperties": false, + "required": [ + "kind", + "uri", + "sha256" + ], + "properties": { + "kind": { + "$ref": "#/$defs/identifier" + }, + "uri": { + "$ref": "#/$defs/nonempty" + }, + "sha256": { + "$ref": "#/$defs/digest" + } + } + }, + "provenance": { + "type": "object", + "additionalProperties": false, + "required": [ + "created_by", + "created_at", + "sources", + "generation_log_uri", + "human_gatekeeper_role" + ], + "properties": { + "created_by": { + "$ref": "#/$defs/nonempty" + }, + "created_at": { + "type": "string", + "format": "date-time" + }, + "sources": { + "type": "array", + "minItems": 1, + "uniqueItems": true, + "items": { + "$ref": "#/$defs/source" + } + }, + "generation_log_uri": { + "$ref": "#/$defs/nonempty" + }, + "human_gatekeeper_role": { + "$ref": "#/$defs/nonempty" + } + } + } + }, + "allOf": [ + { + "if": { + "properties": { + "capability": { + "properties": { + "capability_id": { + "const": "tenpy.finite_1d.dmrg" + } + }, + "required": [ + "capability_id" + ] + } + } + }, + "then": { + "properties": { + "problem": { + "properties": { + "task_family": { + "const": "ground_state_1d_finite" + } + } + }, + "capability": { + "properties": { + "maturity": { + "const": "stable" + } + } + }, + "backend_binding": { + "properties": { + "capability_id": { + "const": "tenpy.finite_1d.dmrg" + }, + "request_schema": { + "const": "tn-agent.tenpy.finite-tfim-dmrg.v1" + } + } + }, + "reference": { + "properties": { + "normalization": { + "const": "total" + } + } + }, + "numerics": { + "not": { + "anyOf": [ + { + "required": [ + "min_sweeps" + ] + }, + { + "required": [ + "entropy_tolerance" + ] + } + ] + } + } + } + } + }, + { + "if": { + "properties": { + "capability": { + "properties": { + "capability_id": { + "const": "tenpy.infinite_1d.vumps" + } + }, + "required": [ + "capability_id" + ] + } + } + }, + "then": { + "properties": { + "problem": { + "properties": { + "task_family": { + "const": "ground_state_1d_infinite" + } + } + }, + "capability": { + "properties": { + "maturity": { + "const": "experimental" + } + } + }, + "backend_binding": { + "properties": { + "capability_id": { + "const": "tenpy.infinite_1d.vumps" + }, + "request_schema": { + "const": "tn-agent.tenpy.infinite-xxz-vumps.v1" + } + } + }, + "reference": { + "properties": { + "normalization": { + "const": "per-site" + } + } + }, + "numerics": { + "required": [ + "min_sweeps", + "entropy_tolerance" + ] + } + } + } + } + ] +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/contracts/reason-codes.md b/tracks/agent-kb/solutions/WangTheoPhys/contracts/reason-codes.md new file mode 100644 index 000000000..702a898b6 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/contracts/reason-codes.md @@ -0,0 +1,62 @@ +# Public gate reason codes + +`gate.py` emits one stable `reason_code`. Codes do not contain local paths, +environment values, exception messages, or rejected input. + +## Successful outcomes + +| Code | Meaning | +|---|---| +| `OK` | A document or Library passed validation. | +| `ACCEPTANCE_PASSED` | Evidence matched the experiment and every acceptance condition passed. | + +Evidence documents use `VALIDATOR_PASS` for a successful required validator, +`REPORTED_ONLY` for a bound diagnostic that cannot affect acceptance, and +`BACKEND_LIMITED` only for a limitation preregistered by the route. + +## Document and structure rejection + +| Code | Meaning | +|---|---| +| `DOCUMENT_DUPLICATE_KEY` | Strict JSON found a duplicate object key. | +| `DOCUMENT_NONFINITE` | NaN or Infinity was present. | +| `DOCUMENT_INVALID_JSON` / `DOCUMENT_INVALID_UTF8` | The bytes are not the accepted JSON encoding. | +| `DOCUMENT_IO_ERROR` / `DOCUMENT_UNSAFE_PATH` / `DOCUMENT_TOO_LARGE` | The input was unreadable, not a regular non-symlink file, or over the limit. | +| `DOCUMENT_NOT_CANONICAL` | A value cannot be represented as canonical JSON. | +| `CLI_USAGE_ERROR` | The command line is incomplete or contains an unsupported argument. | +| `INTERNAL_ERROR` | An unexpected implementation failure was converted to a sanitized JSON rejection; no traceback is exposed. | +| `SECURE_FILE_IO_UNAVAILABLE` | The platform cannot provide the required no-follow read primitive. No unsafe compatibility mode is used. | +| `SCHEMA_VERSION_UNSUPPORTED` | The exact version is not supported. | +| `UNKNOWN_FIELD` / `MISSING_FIELD` | The object is lossy or incomplete. | +| `TYPE_MISMATCH` / `VALUE_INVALID` | A field has the wrong JSON type or violates a closed value constraint. | + +## Scientific authority and acceptance rejection + +| Code | Meaning | +|---|---| +| `UNSUPPORTED_ROUTE` | The requested physics/algorithm combination is not one of the two promoted capabilities. | +| `SCIENTIFIC_EVIDENCE_UNATTESTED` | A valid candidate definition lacks a trusted runner receipt or independently checkable state certificate, so it cannot receive scientific acceptance. | +| `BINDING_MISMATCH` | Capability, adapter, backend, or request/result schema differs from the preregistration. | +| `EXPERIMENT_DIGEST_MISMATCH` / `RESULT_DIGEST_MISMATCH` | Canonical content does not match its claimed identity. | +| `EXECUTION_NOT_SUCCEEDED` | Execution is not a successful, non-retryable terminal result. | +| `PROVENANCE_MISMATCH` | Plan, request, generator, raw-result, or normalized-result identities do not form the registered chain. | +| `OBSERVABLE_SET_MISMATCH` / `OBSERVABLE_STATUS_INVALID` | Observable evidence is missing, extra, or incorrectly classified. | +| `VALIDATOR_SET_MISMATCH` / `VALIDATOR_POLICY_MISMATCH` | Validator evidence or preregistration policy differs from the route contract. | +| `VALIDATOR_STATUS_INVALID` / `VALIDATOR_FAILED` | A validator result is internally inconsistent or did not pass. | +| `VALIDATOR_THRESHOLD_FAILED` | A metric violates its preregistered `max`, `min`, or `equals` threshold. | +| `ACCEPTANCE_CONTRACT_INVALID` | Required, reported-only, and backend-limited validator sets do not close exactly. | +| `EVIDENCE_ARTIFACT_MISSING` | Evidence points to a digest that was not verified from disk. | + +## Artifact and Library rejection + +| Code | Meaning | +|---|---| +| `ARTIFACT_IO_ERROR` / `ARTIFACT_UNSAFE_PATH` | An artifact cannot be read safely under the explicit root. | +| `ARTIFACT_TOO_LARGE` / `ARTIFACT_LIMIT_EXCEEDED` | Per-file, count, or aggregate limits were exceeded. | +| `ARTIFACT_DIGEST_MISMATCH` | Raw bytes or size differ from the evidence manifest. | +| `LIBRARY_RECORD_INVALID` / `LIBRARY_RECORD_LIMIT` | A JSONL record is invalid or the record limit was exceeded. | +| `LIBRARY_SEQUENCE_INVALID` | Append order, revision sequence, contradiction, or supersession history is invalid. | + +Exit status `2` is a document/contract rejection. Exit status `3` is evidence +that cannot pass the preregistered contract and its JSON output includes +`"accepted": false`. A successful command exits `0`. diff --git a/tracks/agent-kb/solutions/WangTheoPhys/contracts/validator-evidence-v1.schema.json b/tracks/agent-kb/solutions/WangTheoPhys/contracts/validator-evidence-v1.schema.json new file mode 100644 index 000000000..12120e460 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/contracts/validator-evidence-v1.schema.json @@ -0,0 +1,86 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "https://github.com/QuantumBFS/quantum.harness/blob/main/tracks/agent-kb/solutions/WangTheoPhys/contracts/validator-evidence-v1.schema.json", + "title": "WangTheoPhys Derived Validator Evidence v1", + "description": "Gate-derived required metrics and explicitly reported-only diagnostics, bound to primary, repeat, and preregistered-reference identities.", + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "request_digest", + "backend_result_digest", + "repeat_backend_result_digest", + "reference_artifact_digest", + "results", + "result_digest" + ], + "properties": { + "schema_version": { + "const": "wangtheophys.tn-validator-evidence.v1" + }, + "request_digest": { + "$ref": "#/$defs/digest" + }, + "backend_result_digest": { + "$ref": "#/$defs/digest" + }, + "repeat_backend_result_digest": { + "$ref": "#/$defs/digest" + }, + "reference_artifact_digest": { + "$ref": "#/$defs/digest" + }, + "results": { + "type": "array", + "minItems": 1, + "maxItems": 64, + "items": { + "$ref": "#/$defs/result" + } + }, + "result_digest": { + "$ref": "#/$defs/digest" + } + }, + "$defs": { + "identifier": { + "type": "string", + "pattern": "^[A-Za-z0-9][A-Za-z0-9_.:@-]{0,127}$" + }, + "digest": { + "type": "string", + "pattern": "^sha256:[0-9a-f]{64}$" + }, + "result": { + "type": "object", + "additionalProperties": false, + "required": [ + "id", + "metric", + "value", + "source" + ], + "properties": { + "id": { + "$ref": "#/$defs/identifier" + }, + "metric": { + "type": [ + "string", + "null" + ] + }, + "value": { + "type": [ + "number", + "null" + ] + }, + "source": { + "type": "string", + "minLength": 1 + } + } + } + } +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/docs/plans/2026-07-29-capsule-trust-closure-design.md b/tracks/agent-kb/solutions/WangTheoPhys/docs/plans/2026-07-29-capsule-trust-closure-design.md new file mode 100644 index 000000000..66fc852a0 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/docs/plans/2026-07-29-capsule-trust-closure-design.md @@ -0,0 +1,92 @@ +# Capsule Trust and Route Closure Design + +## Decision + +The current repeat chain is useful evidence that two differently addressed +records were supplied, but it is not trusted evidence that a scheduler ran two +independent attempts. `reproducibility` therefore becomes a +`reported_only` structural diagnostic and is excluded from +`all_required`. A future contract may promote it only when a preregistered +distinct attempt nonce and runner identity are bound to the experiment and +request digests by a trusted scheduler signature/MAC or an external registry +receipt. + +The capsule keeps the repeat delta, exact reconstruction, distinct execution +handle, and distinct raw-result identity checks. These checks detect accidental +reuse and stale artifact chains; they do not establish process independence. +Cosmetic byte differences, warning text, or a second self-authored handle do +not increase scientific assurance. + +## Route closure + +Every experiment accepted by `validate_experiment` must be translatable +without `KeyError` into an exact main-repository TeNPy request and must provide +every observable used by the public gate. Both promoted routes therefore +require `energy` and `variance`; infinite variance remains +`backend_limited`. + +Numerics are route-aware. Finite experiments do not contain `min_sweeps` or +`entropy_tolerance`; the translator records the backend-fixed values +`min_sweeps=0` and `entropy_tolerance=null`. Infinite experiments explicitly +preregister both fields and require +`finite_entanglement_fit.max_chi == max_bond_dim == chi_schedule[-1]`. + +The infinite Hamiltonian is + +`H = Σ_i [Jxy(Sx_i Sx_{i+1} + Sy_i Sy_{i+1}) + Jxy·Delta Sz_i Sz_{i+1} - h Sz_i]`. + +`Delta` is dimensionless and the external TN-Agent worker implements +`Jz = Jxy·Delta`. That implementation is not part of this QuantumBFS PR's +standalone trust root, so the public capsule exposes only `Jxy=1`. Non-unit +`Jxy` remains fail-closed until an exact worker source identity or trusted +execution receipt is included in the public evidence boundary. + +## Candidate trust boundary + +`candidate` is a valid preregistration status, not a scientific verdict. The +current capsule has no trusted runner receipt and no independently checkable +state/certificate evaluator, so `evaluate()` rejects every candidate with +`SCIENTIFIC_EVIDENCE_UNATTESTED`. Only synthetic `test_fixture` documents may +reach fixture-level `ACCEPTANCE_PASSED`; that outcome proves contract closure, +not fresh tensor-network execution or any success tier of issue #133. + +## Library grounding + +Repository-skill sources resolve only as normalized POSIX paths relative to +the quantum.harness repository root derived from the team directory. Contract +audit evidence resolves only inside the team directory. Validation opens +those files through pinned directory descriptors, rejects traversal, +symlinks, hardlinks, missing files, and content changes, then recomputes the +declared SHA-256. + +Contract-audit evidence carries both `uri` and `sha256`; a digest without an +addressable artifact is not grounded. Append-only sequence rules remain local +consistency rules. Publication still requires an external Git commit/tip or +equivalent immutable registry identity; the file alone cannot prove that +history was never rewritten. + +## Artifact safety + +Every registered artifact must be a stable regular file with exactly one hard +link. The descriptor-based read rejects `st_nlink != 1` before consuming +content and verifies the link count again after reading. + +## Test boundary + +Tests cover: + +- reproduction excluded from required acceptance and retained as + `reported_only`; +- physics-equal repeat records that differ only in whitespace or warning text + remain structural diagnostics, never independent reproduction evidence; +- energy/variance dependency closure and stable rejection instead of + `KeyError`; +- finite backend-fixed `min_sweeps=0`; +- infinite `max_chi == max_bond_dim`; +- non-unit `Jxy` rejected by the standalone capsule while the external worker + retains its independently tested `Jz=Jxy·Delta` implementation; +- fully rebuilt synthetic candidate evidence remaining scientifically + unattested; +- source/evidence traversal, tamper, missing-file, and hardlink rejection; +- standalone behavior plus optional main-repository parser/worker + compatibility. diff --git a/tracks/agent-kb/solutions/WangTheoPhys/docs/plans/2026-07-30-issue133-five-new-problems-design.md b/tracks/agent-kb/solutions/WangTheoPhys/docs/plans/2026-07-30-issue133-five-new-problems-design.md new file mode 100644 index 000000000..e31c09178 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/docs/plans/2026-07-30-issue133-five-new-problems-design.md @@ -0,0 +1,36 @@ +# Issue #133 Five-New-Problem Campaign Design + +## Decision + +Publish five genuinely new, small tensor-network problems inside the public +WangTheoPhys capsule. The historical #124--#128 calibration problems do not +count toward the live campaign. + +Each problem is frozen with an exact-integer input and acceptance rule before +the Solver certificate is emitted. A separate dependency-free Verifier CLI +checks the frozen challenge, the separately materialized gate, and the Solver +certificate in a fresh process. Every positive receipt is paired with a +deterministically corrupted certificate that the same gate must reject. + +## Problems + +1. Minimal MPO rank of a frozen operator matrix. +2. Globally optimal contraction of a frozen four-tensor matrix chain. +3. Exact spectral gap of a frozen transfer matrix. +4. Exact Schmidt rank of a frozen bipartite coefficient matrix. +5. Exact gauge equivalence of two frozen bond-dimension-two MPS tensor sets. + +## Trust boundary + +`human.junkaiwang` is recorded as `human expert supervision` and accepts all +five submission problems and their preregistered gates. The campaign reports +five human-supervised acceptances and five independently executed exact gate +passes. QuantumBFS maintainers retain authority over upstream catalog and +tier determination. No refereed publication is claimed. + +## Public artifacts + +The campaign directory contains Solver and Verifier sources, a deterministic +runner, direct tests, one JSON file for every challenge/gate/certificate/ +negative-control/acceptance/receipt, a campaign manifest, a readable report, +and a SHA-256 file manifest. diff --git a/tracks/agent-kb/solutions/WangTheoPhys/docs/superpowers/plans/2026-07-29-independent-acceptance-evidence.md b/tracks/agent-kb/solutions/WangTheoPhys/docs/superpowers/plans/2026-07-29-independent-acceptance-evidence.md new file mode 100644 index 000000000..3fa264586 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/docs/superpowers/plans/2026-07-29-independent-acceptance-evidence.md @@ -0,0 +1,175 @@ +# Independent Acceptance Evidence Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Remove self-report and scheduler-trust loops while closing every validated route over exact request translation, artifact safety, and grounded Library evidence. + +**Architecture:** The experiment locks an energy reference value, normalization, and source artifact digest. Evidence contains separately addressed primary and repeat records, but repeat consistency remains `reported_only` until a trusted scheduler receipt exists. Route-aware validation guarantees energy/variance dependencies, finite backend-fixed numerics, infinite chi closure, correct XXZ coupling semantics, grounded Library paths, and single-link artifact reads. + +**Tech Stack:** Python 3 standard library, JSON Schema draft 2020-12, `unittest`, optional `jsonschema` compatibility checks. + +## Global Constraints + +- Edit only `tracks/agent-kb/solutions/WangTheoPhys/`. +- Keep `gate.py` dependency-free and fail closed. +- Do not claim numerical or process independence that the public artifacts cannot establish. +- Reproduction cannot enter `all_required` without a preregistered nonce, runner identity, request/experiment digests, and trusted scheduler or registry attestation. +- Direct tests must run in a standalone `quantum.harness` checkout. +- Do not stage, commit, or push. + +--- + +### Task 1: Contract the assurance boundary + +**Files:** +- Modify: `gate.py` +- Modify: `contracts/experiment-v1.schema.json` +- Modify: `contracts/evidence-v1.schema.json` +- Modify: `contracts/validator-evidence-v1.schema.json` +- Create: `contracts/energy-reference-v1.schema.json` +- Test: `tests/test_gate.py` + +**Interfaces:** +- Consumes: experiment `reference`, validator `policy`, and two execution records. +- Produces: exact `required_pass`, `reported_only`, and `backend_limited` policy sets. + +- [x] Add a failing test that rejects a finite experiment whose reported-only validator is placed in `required_validator_ids`. +- [x] Run the direct test and confirm `reported_only` is unsupported before implementation. +- [x] Add strict reference/source validation and exact validator-policy closure. +- [x] Run the direct policy tests and confirm they pass. + +### Task 2: Validate two distinct artifact chains + +**Files:** +- Modify: `gate.py` +- Modify: `contracts/evidence-v1.schema.json` +- Test: `tests/test_gate.py` + +**Interfaces:** +- Consumes: primary and repeat execution evidence plus exact artifact roles. +- Produces: reconstructed primary/repeat backend bundles and their semantic digests. + +- [x] Add failing tests for a missing repeat artifact, reused execution handle, and reused raw artifact identity. +- [x] Run those tests and confirm the old six-artifact contract fails the new expectations. +- [x] Validate exact primary/repeat stream bindings, rebuild both normalized bundles, and reject identical handles/raw identities. +- [x] Run the artifact-chain tests and confirm they pass. + +### Task 3: Derive acceptance without self-report loops + +**Files:** +- Modify: `gate.py` +- Modify: `contracts/validator-evidence-v1.schema.json` +- Test: `tests/test_gate.py` + +**Interfaces:** +- Consumes: primary energy, repeat energy, preregistered reference record, and reported raw diagnostics. +- Produces: benchmark and reproduction deltas that cannot be supplied by the primary raw record. + +- [x] Add a coherent full-rebuild `energy=999` attack test that rewrites every downstream digest. +- [x] Confirm the old gate accepts the coherent attack. +- [x] Compute benchmark delta only from the locked reference and reproduction delta only from the repeat raw result; label raw-only diagnostics `reported_only`. +- [x] Confirm the coherent attack fails with `VALIDATOR_THRESHOLD_FAILED`. + +### Task 4: Regenerate fixtures and public documentation + +**Files:** +- Modify: `fixtures/valid-finite/**` +- Modify: `fixtures/valid-infinite/**` +- Modify: `README.md` +- Modify: `skill/tn-agent-workflow/SKILL.md` +- Modify: `contracts/reason-codes.md` +- Test: `tests/test_gate.py` + +**Interfaces:** +- Consumes: finite exact reference `-8.749171017567908` and infinite analytic reference `0.25-ln(2)`. +- Produces: internally closed synthetic contract fixtures with explicit non-scientific status. + +- [x] Add reference records and distinct repeat raw/result/stream artifacts for both fixtures. +- [x] Regenerate every semantic and byte digest from the final files. +- [x] Remove all wording that calls reported diagnostics fresh or independent. +- [x] Run direct evaluation for both fixtures and confirm `ACCEPTANCE_PASSED`. + +### Task 5: Standalone and repository gates + +**Files:** +- Modify: `tests/test_gate.py` + +**Interfaces:** +- Consumes: optional discoverable TN-Agent checkout and optional `jsonschema`. +- Produces: standalone direct tests with compatibility checks skipped only when dependencies are absent. + +- [x] Replace hard-coded sibling `.venv` assertions with optional discovery and `skipTest`. +- [x] Run `python3 -m unittest discover -s tracks/agent-kb/solutions/WangTheoPhys/tests -v`. +- [x] Run capsule `pytest`, repository `scripts` tests, Ruff format/check, schema validation, digest validation, and `git diff --check`. +- [x] Record remaining limitations: preregistration supplies an anchor, while the gate does not independently rerun the solver or verify MPS state certificates. + +### Task 6: Remove scheduler trust from acceptance + +**Files:** +- Modify: `gate.py` +- Modify: `fixtures/regenerate.py` +- Modify: `README.md` +- Modify: `skill/tn-agent-workflow/SKILL.md` +- Test: `tests/test_gate.py` + +**Interfaces:** +- Consumes: separately addressed repeat raw/result/execution records. +- Produces: a `reported_only` reproduction delta that never affects `ACCEPTANCE_PASSED`. + +- [x] Make `reproducibility` part of the route's reported-only validator set with null operator/threshold. +- [x] Remove it from every fixture's `required_validator_ids`. +- [x] Test whitespace/warning-only repeat differences and coherent repeat rebuilds without describing them as independent execution. +- [x] Document the exact trusted receipt fields required for future promotion. + +### Task 7: Close exact route translation + +**Files:** +- Modify: `gate.py` +- Modify: `contracts/experiment-v1.schema.json` +- Modify: `fixtures/valid-finite/experiment.json` +- Modify: `fixtures/regenerate.py` +- Test: `tests/test_gate.py` + +**Interfaces:** +- Consumes: validated finite/infinite experiment documents. +- Produces: total `_expected_request` translation accepted by the main parser and worker validators. + +- [x] Require `energy` and `variance` in every promoted experiment and JSON Schema route. +- [x] Remove finite experiment `min_sweeps` and `entropy_tolerance`; translate backend-fixed values to `0` and `null`. +- [x] Require infinite `fit.max_chi == max_bond_dim == chi_schedule[-1]`. +- [x] Parameterize exact route closure and verify `Jz=Jxy*Delta` in the + optional sibling-worker compatibility probe. The superseding release + gate restricts this standalone capsule to `Jxy=1` until that worker is + inside the public trust root. + +### Task 8: Ground Library identities + +**Files:** +- Modify: `gate.py` +- Modify: `library/heuristic-v1.schema.json` +- Modify: `library/heuristics.jsonl` +- Modify: `library/README.md` +- Test: `tests/test_gate.py` + +**Interfaces:** +- Consumes: repository-skill and contract-audit `uri`/SHA-256 pairs. +- Produces: descriptor-verified, path-confined Library records. + +- [x] Add `uri` and `sha256` to contract-audit evidence. +- [x] Resolve repository-skill paths only relative to the repository root and contract-audit paths only relative to the team root. +- [x] Recompute identities and reject traversal, missing, hardlinked, symlinked, or tampered sources. +- [x] Document that append-only history needs an externally frozen Git tip. + +### Task 9: Reject artifact hardlinks + +**Files:** +- Modify: `gate.py` +- Test: `tests/test_gate.py` + +**Interfaces:** +- Consumes: descriptor-opened registered artifacts. +- Produces: rejection when either pre-read or post-read `st_nlink != 1`. + +- [x] Add a hardlinked artifact regression test. +- [x] Reject multi-link files before reading and after the stability check. +- [x] Run standalone, monorepo, schema, determinism, attack, repository scripts, Ruff, and diff gates. diff --git a/tracks/agent-kb/solutions/WangTheoPhys/docs/superpowers/plans/2026-07-29-release-blocker-closure.md b/tracks/agent-kb/solutions/WangTheoPhys/docs/superpowers/plans/2026-07-29-release-blocker-closure.md new file mode 100644 index 000000000..f6f61ea11 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/docs/superpowers/plans/2026-07-29-release-blocker-closure.md @@ -0,0 +1,263 @@ +# QuantumBFS Capsule Release Blocker Closure Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Close the final three independent-audit blockers so the WangTheoPhys capsule rejects unattested candidates, exposes only standalone-proven XXZ semantics, and binds Library record kinds to exact path classes. + +**Architecture:** Keep experiment validation separate from scientific acceptance: `candidate` definitions remain valid preregistrations, but `evaluate()` rejects them until a trusted runner receipt or independently checkable state certificate exists. Restrict the public infinite route to `Jxy == 1.0` because the external worker is outside this PR's trust root. Validate Library kind/path pairs before descriptor-based content verification, then keep the existing URI confinement, file identity, and SHA-256 checks. + +**Tech Stack:** Python 3 standard library, JSON Schema draft 2020-12, `unittest`, optional `jsonschema`, Ruff. + +## Global Constraints + +- Edit only `tracks/agent-kb/solutions/WangTheoPhys/`. +- Keep `gate.py` dependency-free, deterministic, bounded, and fail closed. +- Preserve `candidate` as a valid preregistration status, but never return `accepted: true` for it without a future attestation contract. +- A scientific rejection uses exit status `3` and emits `accepted: false`. +- The standalone capsule supports infinite XXZ only when `Jxy == 1.0`. +- Repository skills and method/workflow cards use `skills//SKILL.md`; contract audits use a regular file below `docs/` or `tests/` in the team directory. +- Run every direct test from a standalone `quantum.harness` checkout; optional TN-Agent integration may skip only when the sibling dependency is absent. +- Do not stage, commit, push, or update PR #209 until every release gate in Task 4 passes. + +--- + +### Task 1: Reject unattested candidate acceptance + +**Files:** +- Modify: `gate.py` +- Modify: `contracts/reason-codes.md` +- Modify: `README.md` +- Modify: `skill/tn-agent-workflow/SKILL.md` +- Test: `tests/test_gate.py` + +**Interfaces:** +- Consumes: `validate_experiment(document) -> dict[str, object]` and its `problem_status` field. +- Produces: `evaluate(...)` rejection code `SCIENTIFIC_EVIDENCE_UNATTESTED`, exit status `3`, and JSON field `accepted: false` for a validated `candidate`. + +- [ ] **Step 1: Write the synthetic-candidate attack test** + +```python +def test_synthetic_candidate_cannot_self_report_scientific_acceptance(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + shutil.copytree(self.fixture("valid-finite"), root / "fixtures/valid-finite") + experiment_path = root / "fixtures/valid-finite/experiment.json" + experiment = json.loads(experiment_path.read_text(encoding="utf-8")) + experiment["problem"]["status"] = "candidate" + experiment_path.write_text(json.dumps(experiment), encoding="utf-8") + regenerate = load_regenerator() + regenerate.SOLUTION_ROOT = root + regenerate.write_fixture("valid-finite", regenerate.CONFIG["valid-finite"]) + with self.assertRaises(gate.GateError) as caught: + gate.evaluate( + gate.load_json_document(experiment_path), + gate.load_json_document(root / "fixtures/valid-finite/evidence.json"), + artifact_root=root / "fixtures/valid-finite/artifacts", + ) + self.assertEqual(caught.exception.reason_code, "SCIENTIFIC_EVIDENCE_UNATTESTED") + self.assertEqual(caught.exception.exit_code, 3) + self.assertEqual(caught.exception.as_dict()["accepted"], False) +``` + +- [ ] **Step 2: Run the attack test and verify the current gate accepts the rebuilt candidate** + +Run: `python3 -m unittest discover -s tracks/agent-kb/solutions/WangTheoPhys/tests -p 'test_gate.py' -k test_synthetic_candidate_cannot_self_report_scientific_acceptance -v` + +Expected before implementation: FAIL because `evaluate()` reaches `ACCEPTANCE_PASSED` or because `SCIENTIFIC_EVIDENCE_UNATTESTED` is unregistered. + +- [ ] **Step 3: Add the explicit trust guard** + +```python +if experiment_summary["problem_status"] == "candidate": + _fail( + "SCIENTIFIC_EVIDENCE_UNATTESTED", + "Candidate evidence lacks a trusted execution or state certificate", + field="$.problem.status", + exit_code=3, + ) +``` + +Add `SCIENTIFIC_EVIDENCE_UNATTESTED` to `GATE_REASON_CODES`. Make `GateError.as_dict()` add `"accepted": False` whenever `exit_code == 3`. + +- [ ] **Step 4: Tighten public claims** + +Document that `ACCEPTANCE_PASSED` means fixture contract closure only, primary energy and convergence remain unattested worker assertions, `candidate` evaluation is rejected, and this PR achieves no success tier of issue #133. + +- [ ] **Step 5: Run candidate, CLI, reason-code, and fixture tests** + +Run: `python3 -m unittest discover -s tracks/agent-kb/solutions/WangTheoPhys/tests -p 'test_gate.py' -k test_synthetic_candidate_cannot_self_report_scientific_acceptance -k test_valid_finite_and_infinite_evaluations -k test_reason_code_document_covers_the_executable_registry -v` + +Expected: PASS; both `test_fixture` evaluations still return `ACCEPTANCE_PASSED`. + +### Task 2: Restrict standalone infinite XXZ to unit Jxy + +**Files:** +- Modify: `gate.py` +- Modify: `contracts/experiment-v1.schema.json` +- Modify: `README.md` +- Modify: `skill/tn-agent-workflow/SKILL.md` +- Modify: `docs/plans/2026-07-29-capsule-trust-closure-design.md` +- Modify: `library/heuristics.jsonl` +- Test: `tests/test_gate.py` + +**Interfaces:** +- Consumes: `physics.model.couplings.Jxy` for `tenpy.infinite_1d.vumps`. +- Produces: runtime `UNSUPPORTED_ROUTE` and JSON Schema rejection for every value other than numeric `1.0`. + +- [ ] **Step 1: Replace the non-unit positive test with a fail-closed test** + +```python +def test_nonunit_jxy_is_outside_the_standalone_capsule_trust_root(self) -> None: + experiment = self.load("valid-infinite/experiment.json") + experiment["physics"]["model"]["couplings"]["Jxy"] = 2.0 + self.assert_reason( + "UNSUPPORTED_ROUTE", lambda: gate.validate_experiment(experiment) + ) +``` + +The optional sibling integration test must assert the fixture's unit-`Jxy` request only; it must not use sibling code to broaden this PR's public route. + +- [ ] **Step 2: Run the non-unit test and confirm the existing route accepts it** + +Run: `python3 -m unittest discover -s tracks/agent-kb/solutions/WangTheoPhys/tests -p 'test_gate.py' -k test_nonunit_jxy_is_outside_the_standalone_capsule_trust_root -v` + +Expected before implementation: FAIL because `validate_experiment()` returns `OK`. + +- [ ] **Step 3: Enforce the runtime and Schema boundary** + +```python +if capability_id == "tenpy.infinite_1d.vumps" and couplings["Jxy"] != 1.0: + _fail( + "UNSUPPORTED_ROUTE", + "Standalone infinite XXZ is limited to Jxy=1", + field="$.physics.model.couplings.Jxy", + ) +``` + +Set the JSON Schema property to `{"type": "number", "const": 1.0}`. Keep the documented Hamiltonian convention `Jz = Jxy * Delta`, but state that this PR exposes only the unit-`Jxy` slice until an attested worker implementation is part of the trust root. + +- [ ] **Step 4: Update the hashed contract-audit record** + +After revising `docs/plans/2026-07-29-capsule-trust-closure-design.md`, recompute its SHA-256 and update only the `contract_audit` evidence digest in `library/heuristics.jsonl`. + +- [ ] **Step 5: Run route, Schema, Library, and fixture tests** + +Run: `python3 -m unittest discover -s tracks/agent-kb/solutions/WangTheoPhys/tests -p 'test_gate.py' -k test_nonunit_jxy_is_outside_the_standalone_capsule_trust_root -k test_public_json_schemas_accept_all_promoted_fixtures -k test_library_is_append_only_and_cross_references_prior_records -v` + +Expected: PASS. + +### Task 3: Bind Library kinds to exact path classes + +**Files:** +- Modify: `gate.py` +- Modify: `library/heuristic-v1.schema.json` +- Modify: `library/README.md` +- Test: `tests/test_gate.py` + +**Interfaces:** +- Consumes: `source.kind`, `source.uri`, `evidence.kind`, and `evidence.uri`. +- Produces: stable `LIBRARY_RECORD_INVALID` before file-content verification when a kind/path pair is semantically invalid. + +- [ ] **Step 1: Add kind/path-confusion attack tests** + +```python +for section, kind, uri, root in ( + ("source", "repository_skill", "README.md", REPOSITORY_ROOT), + ("evidence", "method_card", "README.md", REPOSITORY_ROOT), + ("evidence", "workflow_card", "README.md", REPOSITORY_ROOT), + ("evidence", "contract_audit", "README.md", SOLUTION_ROOT), +): + record = copy.deepcopy(original) + record[section]["kind"] = kind + record[section]["uri"] = uri + record[section]["sha256"] = file_sha256(root / uri) + self.assert_reason("LIBRARY_RECORD_INVALID", lambda: validate_one(record)) +``` + +- [ ] **Step 2: Run the tests and confirm correct-SHA root README files are currently accepted** + +Run: `python3 -m unittest discover -s tracks/agent-kb/solutions/WangTheoPhys/tests -p 'test_gate.py' -k test_library_kind_and_path_must_match -v` + +Expected before implementation: FAIL because `validate_library()` returns `OK` for at least one mutation. + +- [ ] **Step 3: Implement a single runtime kind/path validator** + +```python +SKILL_URI_PATTERN = re.compile(r"^skills/[a-z0-9][a-z0-9-]{0,63}/SKILL[.]md$") +AUDIT_URI_PATTERN = re.compile(r"^(?:docs|tests)/(?:[A-Za-z0-9._-]+/)*[A-Za-z0-9._-]+$") + + +def _validate_library_kind_uri(kind: str, uri: str, field: str) -> None: + pattern = ( + SKILL_URI_PATTERN + if kind in {"repository_skill", "method_card", "workflow_card"} + else AUDIT_URI_PATTERN + ) + if pattern.fullmatch(uri) is None: + _fail( + "LIBRARY_RECORD_INVALID", + "Library kind and path do not match", + field=f"{field}.uri", + ) +``` + +Call this before `_validate_grounded_library_file()` for both source and evidence. + +- [ ] **Step 4: Mirror the constraint in JSON Schema** + +Make `source.kind` a constant `repository_skill`; constrain source URI to the skill pattern. Make `evidence.kind` an enum and use `if/then` branches so method/workflow cards use the skill pattern and contract audits use the `docs/`/`tests/` pattern. + +- [ ] **Step 5: Run Library positive and negative tests** + +Run: `python3 -m unittest discover -s tracks/agent-kb/solutions/WangTheoPhys/tests -p 'test_gate.py' -k test_library_is_append_only_and_cross_references_prior_records -k test_library_grounding_rejects_tamper_traversal_and_missing_files -k test_library_kind_and_path_must_match -v` + +Expected: PASS. + +### Task 4: Execute the release gate + +**Files:** +- Modify only if a reproducible gate failure proves a scoped defect in Tasks 1-3. + +**Interfaces:** +- Consumes: the complete WangTheoPhys public capsule. +- Produces: reproducible evidence that the capsule is ready to commit and update in PR #209. + +- [ ] **Step 1: Run direct capsule unittest** + +Run: `python3 -m unittest discover -s tracks/agent-kb/solutions/WangTheoPhys/tests -v` + +Expected: all tests pass; only explicitly optional integration dependencies may skip. + +- [ ] **Step 2: Run capsule pytest with the TN-Agent integration environment** + +Run: `TN_AGENT_STARTER_ROOT=/Users/thomasjwang/Documents/GitHub/Projects/Agents/Tensor_Network/tn-agent-starter TN_AGENT_INTEGRATION_PYTHON=/Users/thomasjwang/Documents/GitHub/Projects/Agents/Tensor_Network/tn-agent-starter/.venv/bin/python python3 -m pytest tracks/agent-kb/solutions/WangTheoPhys/tests -q` + +Expected: all capsule tests and subtests pass. + +- [ ] **Step 3: Prove deterministic fixture regeneration** + +Copy the capsule to two fresh temporary directories, run `fixtures/regenerate.py` in each, and compare every regular file byte-for-byte. The checked-in capsule must equal a fresh regeneration except for the implementation-plan documents, which are not generated artifacts. + +- [ ] **Step 4: Simulate a standalone checkout** + +Copy the repository without the sibling TN-Agent checkout, run direct unittest, and confirm only optional TN-Agent integration checks skip; candidate, unit-`Jxy`, and Library confusion tests must still execute and pass. + +- [ ] **Step 5: Run the full harness suite and static gates** + +Run: `python3 -m pytest -q` + +Run: `ruff check tracks/agent-kb/solutions/WangTheoPhys` + +Run: `ruff format --check tracks/agent-kb/solutions/WangTheoPhys` + +Run: `git diff --check` + +Expected: all pass. + +- [ ] **Step 6: Inspect the exact diff and commit only the capsule** + +Stage only `tracks/agent-kb/solutions/WangTheoPhys/`, verify `git diff --cached --check`, and commit with message `Close WangTheoPhys capsule trust boundaries`. + +- [ ] **Step 7: Update existing PR #209 only after the commit and push succeed** + +Push `challenge/agent-wangtheophys`, confirm PR #209 points to the pushed commit, and state explicitly in the PR description: fixture contract closure only, no fresh trusted solve, no issue #133 success tier, and `candidate` evidence remains scientifically unattested. diff --git a/tracks/agent-kb/solutions/WangTheoPhys/docs/superpowers/plans/2026-07-30-issue133-five-new-problems.md b/tracks/agent-kb/solutions/WangTheoPhys/docs/superpowers/plans/2026-07-30-issue133-five-new-problems.md new file mode 100644 index 000000000..a2ba4326c --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/docs/superpowers/plans/2026-07-30-issue133-five-new-problems.md @@ -0,0 +1,76 @@ +# Issue #133 Five-New-Problem Campaign Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Publish five new human-supervised problems and five exact solved-gate receipts in the public #209 capsule. + +**Architecture:** A deterministic Solver emits certificates but never verdicts. A separate standard-library Verifier CLI derives every verdict from frozen challenge and gate documents in fresh subprocesses; the runner publishes all bindings, negative controls, decisions, receipts, and checksums. + +**Tech Stack:** Python 3 standard library, JSON, SHA-256, `unittest`. + +## Global Constraints + +- Edit only `tracks/agent-kb/solutions/WangTheoPhys/`. +- Count five new problems; never count #124--#128 calibration items. +- Use exact integer/rational arithmetic and fail closed. +- Record `human.junkaiwang` as `human expert supervision`. +- Leave refereed publications at `0` and upstream catalog determination pending. + +--- + +### Task 1: Freeze and solve five new problems + +**Files:** +- Create: `issue133-campaign/campaign_solver.py` +- Test: `issue133-campaign/tests/test_campaign.py` + +**Interfaces:** +- Produces: `frozen_challenges()`, `solve_challenge()`, and `negative_control()`. + +- [ ] Define five immutable challenge records and canonical digests. +- [ ] Emit deterministic exact certificates without acceptance claims. +- [ ] Emit one essential corruption for every certificate type. + +### Task 2: Independently verify every gate + +**Files:** +- Create: `issue133-campaign/campaign_verifier.py` +- Test: `issue133-campaign/tests/test_campaign.py` + +**Interfaces:** +- Consumes: one challenge, one separately frozen gate, and one certificate. +- Produces: an exact derived acceptance record or exit code `3`. + +- [ ] Check all challenge, gate, and certificate identity bindings. +- [ ] Derive rank, global contraction optimum, spectrum, and gauge equations. +- [ ] Prove all five positives pass and all five corruptions fail. + +### Task 3: Publish the evidence graph + +**Files:** +- Create: `issue133-campaign/run_campaign.py` +- Create: `issue133-campaign/README.md` +- Generate: `issue133-campaign/artifacts/**` + +**Interfaces:** +- Consumes: Solver and Verifier source identities. +- Produces: five challenge/gate/certificate/negative/acceptance/receipt sets, `campaign.json`, `REPORT.md`, and `SHA256SUMS.txt`. + +- [ ] Materialize all challenges and gates before solving. +- [ ] Execute each positive and negative gate in a fresh subprocess. +- [ ] Bind the authorized human decision and exact verifier result. +- [ ] Generate the campaign manifest, table, and checksums deterministically. + +### Task 4: Integrate and deliver PR #209 + +**Files:** +- Modify: `README.md` +- Modify: PR #209 body + +**Interfaces:** +- Consumes: public campaign manifest and receipt digests. +- Produces: a five-row acceptance/solve table and one-command replay surface. + +- [ ] Run direct tests, replay generation, checksum verification, Ruff, and diff checks. +- [ ] Commit and push the submission branch. +- [ ] Update PR #209 with five rows and exact replay commands. diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/invalid/duplicate-key.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/invalid/duplicate-key.json new file mode 100644 index 000000000..96494da5b --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/invalid/duplicate-key.json @@ -0,0 +1 @@ +{"schema_version":"wangtheophys.tn-experiment.v1","schema_version":"duplicated"} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/invalid/missing-field.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/invalid/missing-field.json new file mode 100644 index 000000000..0fa640bf7 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/invalid/missing-field.json @@ -0,0 +1 @@ +{"schema_version":"wangtheophys.tn-experiment.v1"} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/invalid/nonfinite.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/invalid/nonfinite.json new file mode 100644 index 000000000..48c24bfeb --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/invalid/nonfinite.json @@ -0,0 +1 @@ +{"schema_version":"wangtheophys.tn-experiment.v1","value":NaN} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/invalid/unknown-field.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/invalid/unknown-field.json new file mode 100644 index 000000000..8e5d0407b --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/invalid/unknown-field.json @@ -0,0 +1 @@ +{"schema_version":"wangtheophys.tn-experiment.v1","unexpected":true} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/regenerate.py b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/regenerate.py new file mode 100644 index 000000000..f94859501 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/regenerate.py @@ -0,0 +1,450 @@ +"""Regenerate the synthetic public contract fixtures deterministically.""" + +from __future__ import annotations + +import copy +import hashlib +import json +import sys +from pathlib import Path + +SOLUTION_ROOT = Path(__file__).resolve().parents[1] +sys.path.insert(0, str(SOLUTION_ROOT)) + +import gate + +CONFIG = { + "valid-finite": { + "reference_id": "finite-tfim-l8-open-j1-g0.8-ed", + "reference": -8.749171017567908, + "primary": -8.749171017567908, + "previous": -8.749171007567907, + "repeat": -8.749170997567908, + "repeat_previous": -8.749170987567907, + "normalization": "total", + "method": "exact_diagonalization", + "citation": ( + "Exact diagonalization anchor for the preregistered open L=8 " + "TFIM Hamiltonian with J=1 and g=0.8." + ), + "primary_handle": "fixture:finite:primary", + "repeat_handle": "fixture:finite:repeat", + }, + "valid-infinite": { + "reference_id": "infinite-heisenberg-delta1-bethe-ansatz", + "reference": -0.4431471805599453, + "primary": -0.4431, + "previous": -0.443, + "repeat": -0.44311, + "repeat_previous": -0.44301, + "normalization": "per-site", + "method": "analytic_bethe_ansatz", + "citation": ( + "Thermodynamic-limit spin-1/2 antiferromagnetic Heisenberg-chain " + "ground-state energy per site, 1/4-ln(2)." + ), + "primary_handle": "fixture:infinite:primary", + "repeat_handle": "fixture:infinite:repeat", + }, +} + + +def encoded(value: object) -> bytes: + return ( + json.dumps( + value, + allow_nan=False, + ensure_ascii=False, + indent=2, + ) + + "\n" + ).encode("utf-8") + + +def file_digest(raw: bytes) -> str: + return "sha256:" + hashlib.sha256(raw).hexdigest() + + +def semantic(value: dict[str, object]) -> dict[str, object]: + content = {key: item for key, item in value.items() if key != "result_digest"} + return {**content, "result_digest": gate.canonical_digest(content)} + + +def artifact( + *, + relative_path: str, + role: str, + media_type: str, + raw: bytes, +) -> dict[str, object]: + return { + "relative_path": relative_path, + "digest": file_digest(raw), + "size_bytes": len(raw), + "media_type": media_type, + "role": role, + } + + +def set_policy(experiment: dict[str, object]) -> None: + validators = experiment["validators"] + assert isinstance(validators, list) + for validator in validators: + assert isinstance(validator, dict) + identifier = validator["id"] + if identifier in { + "canonical_form", + "symmetry_check", + "reproducibility", + } or ( + identifier == "variance" + and experiment["backend_binding"]["capability_id"] # type: ignore[index] + == "tenpy.finite_1d.dmrg" + ): + validator["policy"] = "reported_only" + validator["operator"] = None + validator["threshold"] = None + required = [ + str(validator["id"]) + for validator in validators + if isinstance(validator, dict) and validator["policy"] == "required_pass" + ] + limited = [ + str(validator["id"]) + for validator in validators + if isinstance(validator, dict) and validator["policy"] == "backend_limited" + ] + reported = [ + str(validator["id"]) + for validator in validators + if isinstance(validator, dict) and validator["policy"] == "reported_only" + ] + acceptance = experiment["acceptance"] + assert isinstance(acceptance, dict) + acceptance["required_validator_ids"] = required + acceptance["allowed_backend_limited_ids"] = limited + acceptance["reported_only_validator_ids"] = reported + + +def make_raw( + template: dict[str, object], + *, + request_digest: str, + plan_id: str, + energy: float, + previous_energy: float, +) -> dict[str, object]: + raw = copy.deepcopy(template) + raw["request_digest"] = request_digest + raw["plan_id"] = plan_id + observables = raw["observables"] + convergence = raw["convergence"] + assert isinstance(observables, dict) + assert isinstance(convergence, list) + energy_observable = observables["energy"] + assert isinstance(energy_observable, dict) + energy_observable["value"] = energy + previous = convergence[-2] + latest = convergence[-1] + assert isinstance(previous, dict) + assert isinstance(latest, dict) + previous_metrics = previous["metrics"] + latest_metrics = latest["metrics"] + assert isinstance(previous_metrics, dict) + assert isinstance(latest_metrics, dict) + previous_metrics["energy"] = previous_energy + latest["metrics"] = { + "energy": energy, + "canonical_residual": latest_metrics["canonical_residual"], + "symmetry_residual": latest_metrics["symmetry_residual"], + } + return raw + + +def write_fixture(name: str, config: dict[str, object]) -> None: + fixture_root = SOLUTION_ROOT / "fixtures" / name + artifact_root = fixture_root / "artifacts" + experiment = json.loads((fixture_root / "experiment.json").read_text()) + raw_template = json.loads((artifact_root / "backend-raw-result.json").read_text()) + binding = experiment["backend_binding"] + numerics = experiment["numerics"] + assert isinstance(binding, dict) + assert isinstance(numerics, dict) + if binding["capability_id"] == "tenpy.finite_1d.dmrg": + experiment["numerics"] = { + key: value + for key, value in numerics.items() + if key not in {"min_sweeps", "entropy_tolerance"} + } + set_policy(experiment) + + reference_record = semantic( + { + "schema_version": "wangtheophys.tn-energy-reference.v1", + "reference_id": config["reference_id"], + "capability_id": experiment["backend_binding"]["capability_id"], + "physics_digest": gate.canonical_digest(experiment["physics"]), + "observable": "energy", + "value": config["reference"], + "units": "J", + "normalization": config["normalization"], + "method": config["method"], + "citation": config["citation"], + } + ) + reference_raw = encoded(reference_record) + experiment["reference"] = { + "observable": "energy", + "value": config["reference"], + "units": "J", + "normalization": config["normalization"], + "source": { + "kind": "registered_artifact", + "uri": "energy-reference.json", + "sha256": file_digest(reference_raw), + }, + } + experiment_raw = encoded(experiment) + (fixture_root / "experiment.json").write_bytes(experiment_raw) + + experiment_digest = gate.canonical_digest(experiment) + plan_id = gate._expected_plan_id(experiment, experiment_digest) + request = gate._expected_request(experiment, experiment_digest, plan_id) + request_raw = encoded(request) + request_digest = gate.canonical_digest(request) + + primary_raw_value = make_raw( + raw_template, + request_digest=request_digest, + plan_id=plan_id, + energy=float(config["primary"]), + previous_energy=float(config["previous"]), + ) + repeat_raw_value = make_raw( + raw_template, + request_digest=request_digest, + plan_id=plan_id, + energy=float(config["repeat"]), + previous_energy=float(config["repeat_previous"]), + ) + primary_raw = encoded(primary_raw_value) + repeat_raw = encoded(repeat_raw_value) + primary_raw_artifact = artifact( + relative_path="backend-raw-result.json", + role="backend_raw_result", + media_type="application/json", + raw=primary_raw, + ) + repeat_raw_artifact = artifact( + relative_path="backend-repeat-raw-result.json", + role="backend_repeat_raw_result", + media_type="application/json", + raw=repeat_raw, + ) + + primary_stdout = f"{name} primary ok\n".encode() + primary_stderr = f"{name} primary log\n".encode() + repeat_stdout = f"{name} repeat ok\n".encode() + repeat_stderr = f"{name} repeat log\n".encode() + execution = { + "schema_version": "tn-agent.execution-evidence.v1", + "status": "succeeded", + "return_code": 0, + "execution_handle": config["primary_handle"], + "retryable": False, + "stdout_digest": file_digest(primary_stdout), + "stderr_digest": file_digest(primary_stderr), + "stdout_truncated": False, + "stderr_truncated": False, + } + repeat_execution = { + "schema_version": "tn-agent.execution-evidence.v1", + "status": "succeeded", + "return_code": 0, + "execution_handle": config["repeat_handle"], + "retryable": False, + "stdout_digest": file_digest(repeat_stdout), + "stderr_digest": file_digest(repeat_stderr), + "stdout_truncated": False, + "stderr_truncated": False, + } + binding = experiment["backend_binding"] + assert isinstance(binding, dict) + primary_bundle = gate._reconstruct_backend_bundle( + request_digest=request_digest, + binding=binding, + raw=primary_raw_value, + execution=execution, + raw_artifact=primary_raw_artifact, + ) + repeat_bundle = gate._reconstruct_backend_bundle( + request_digest=request_digest, + binding=binding, + raw=repeat_raw_value, + execution=repeat_execution, + raw_artifact=repeat_raw_artifact, + ) + primary_bundle_raw = encoded(primary_bundle) + repeat_bundle_raw = encoded(repeat_bundle) + primary_bundle_artifact = artifact( + relative_path="backend-result.json", + role="backend_result", + media_type="application/json", + raw=primary_bundle_raw, + ) + repeat_bundle_artifact = artifact( + relative_path="backend-repeat-result.json", + role="backend_repeat_result", + media_type="application/json", + raw=repeat_bundle_raw, + ) + + validator_results = gate._derived_validator_results( + primary_raw_value, + repeat_raw=repeat_raw_value, + reference=reference_record, + ) + validator_evidence = semantic( + { + "schema_version": "wangtheophys.tn-validator-evidence.v1", + "request_digest": request_digest, + "backend_result_digest": primary_bundle["result_digest"], + "repeat_backend_result_digest": repeat_bundle["result_digest"], + "reference_artifact_digest": file_digest(reference_raw), + "results": validator_results, + } + ) + validator_raw = encoded(validator_evidence) + validator_artifact = artifact( + relative_path="validator-evidence.json", + role="validator_evidence", + media_type="application/json", + raw=validator_raw, + ) + + artifacts = [ + artifact( + relative_path="backend-request.json", + role="backend_request", + media_type="application/json", + raw=request_raw, + ), + primary_raw_artifact, + primary_bundle_artifact, + repeat_raw_artifact, + repeat_bundle_artifact, + artifact( + relative_path="energy-reference.json", + role="energy_reference", + media_type="application/json", + raw=reference_raw, + ), + validator_artifact, + artifact( + relative_path="backend-stdout.log", + role="backend_stdout", + media_type="text/plain", + raw=primary_stdout, + ), + artifact( + relative_path="backend-stderr.log", + role="backend_stderr", + media_type="text/plain", + raw=primary_stderr, + ), + artifact( + relative_path="backend-repeat-stdout.log", + role="backend_repeat_stdout", + media_type="text/plain", + raw=repeat_stdout, + ), + artifact( + relative_path="backend-repeat-stderr.log", + role="backend_repeat_stderr", + media_type="text/plain", + raw=repeat_stderr, + ), + ] + observable_statuses = gate.ROUTES[str(binding["capability_id"])][ + "observable_statuses" + ] + assert isinstance(observable_statuses, dict) + validators = experiment["validators"] + assert isinstance(validators, list) + metrics = {str(item["id"]): item["value"] for item in validator_results} + public_validator_results: list[dict[str, object]] = [] + for validator in validators: + assert isinstance(validator, dict) + policy = validator["policy"] + if policy == "required_pass": + status = "pass" + reason_code = "VALIDATOR_PASS" + elif policy == "reported_only": + status = "reported_only" + reason_code = "REPORTED_ONLY" + else: + status = "backend_limited" + reason_code = "BACKEND_LIMITED" + public_validator_results.append( + { + "id": validator["id"], + "status": status, + "reason_code": reason_code, + "metric_value": metrics[str(validator["id"])], + "evidence_digest": validator_artifact["digest"], + } + ) + evidence = semantic( + { + "schema_version": "wangtheophys.tn-evidence.v1", + "experiment_digest": experiment_digest, + "binding": binding, + "execution": execution, + "repeat_execution": repeat_execution, + "artifacts": artifacts, + "observables": [ + { + "name": name, + "status": observable_statuses[name], + "evidence_digest": primary_bundle_artifact["digest"], + } + for name in experiment["observables"] + ], + "validator_results": public_validator_results, + "provenance": { + "plan_id": plan_id, + "request_digest": request_digest, + "backend_result_digest": primary_bundle_artifact["digest"], + "repeat_backend_result_digest": repeat_bundle_artifact["digest"], + "generated_by": "fixtures/regenerate.py", + "generated_at": "2026-07-29T00:01:00Z", + }, + } + ) + + files = { + "backend-request.json": request_raw, + "backend-raw-result.json": primary_raw, + "backend-result.json": primary_bundle_raw, + "backend-repeat-raw-result.json": repeat_raw, + "backend-repeat-result.json": repeat_bundle_raw, + "energy-reference.json": reference_raw, + "validator-evidence.json": validator_raw, + "backend-stdout.log": primary_stdout, + "backend-stderr.log": primary_stderr, + "backend-repeat-stdout.log": repeat_stdout, + "backend-repeat-stderr.log": repeat_stderr, + } + for relative_path, raw in files.items(): + (artifact_root / relative_path).write_bytes(raw) + (fixture_root / "evidence.json").write_bytes(encoded(evidence)) + + +def main() -> int: + for name, config in CONFIG.items(): + write_fixture(name, config) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-raw-result.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-raw-result.json new file mode 100644 index 000000000..6b186b24a --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-raw-result.json @@ -0,0 +1,90 @@ +{ + "schema_version": "tn-agent.tenpy.raw-result.v1", + "request_digest": "sha256:247d02d04c82677698405851fe340c440a23b01620870591de0b785a60847a5d", + "plan_id": "sha256:56217e253c4f5e13e66b9939c8ff5f56d26ee77023869288be54c2fe2c0ae307", + "capability_id": "tenpy.finite_1d.dmrg", + "backend_id": "tenpy", + "backend_version": "1.1.0", + "adapter_id": "tenpy.v1", + "adapter_version": "1.0.0", + "status": "succeeded", + "environment": { + "schema_version": "tn-agent.tenpy.environment.v1", + "environment_digest": "sha256:33f97f66d496d52d7a68f6a05160f07f90b1d928cea8481f892db7dac132a756", + "runtime_id": "python", + "runtime_version": "3.12.4", + "dependencies": { + "numpy": "2.0.1", + "tenpy": "1.1.0" + } + }, + "observables": { + "energy": { + "status": "measured", + "value": -8.749171017567908, + "units": "J", + "normalization": "total", + "reason": null + }, + "variance": { + "status": "measured", + "value": 1e-10, + "units": "J^2", + "normalization": "total", + "reason": null + }, + "magnetization_z": { + "status": "measured", + "value": 0.24931, + "units": "dimensionless", + "normalization": "per-site", + "reason": null + }, + "entanglement_entropy": { + "status": "measured", + "value": [ + 0.132, + 0.284, + 0.315, + 0.284, + 0.132 + ], + "units": "dimensionless", + "normalization": "bond", + "reason": null + }, + "correlator_sigma_x": { + "status": "measured", + "value": [ + 1.0, + 0.731, + 0.548, + 0.402 + ], + "units": "dimensionless", + "normalization": "distance", + "reason": null + } + }, + "convergence": [ + { + "step": 1, + "metrics": { + "energy": -8.749171007567908 + } + }, + { + "step": 2, + "metrics": { + "energy": -8.749171017567908, + "canonical_residual": 1e-10, + "symmetry_residual": 1e-12 + } + } + ], + "warnings": [], + "known_limitations": [ + "Only the finite spin-1/2 TFIM open-chain route is promoted.", + "The fixture does not count as a generated challenge problem." + ] +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-repeat-raw-result.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-repeat-raw-result.json new file mode 100644 index 000000000..a1e083847 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-repeat-raw-result.json @@ -0,0 +1,90 @@ +{ + "schema_version": "tn-agent.tenpy.raw-result.v1", + "request_digest": "sha256:247d02d04c82677698405851fe340c440a23b01620870591de0b785a60847a5d", + "plan_id": "sha256:56217e253c4f5e13e66b9939c8ff5f56d26ee77023869288be54c2fe2c0ae307", + "capability_id": "tenpy.finite_1d.dmrg", + "backend_id": "tenpy", + "backend_version": "1.1.0", + "adapter_id": "tenpy.v1", + "adapter_version": "1.0.0", + "status": "succeeded", + "environment": { + "schema_version": "tn-agent.tenpy.environment.v1", + "environment_digest": "sha256:33f97f66d496d52d7a68f6a05160f07f90b1d928cea8481f892db7dac132a756", + "runtime_id": "python", + "runtime_version": "3.12.4", + "dependencies": { + "numpy": "2.0.1", + "tenpy": "1.1.0" + } + }, + "observables": { + "energy": { + "status": "measured", + "value": -8.749170997567909, + "units": "J", + "normalization": "total", + "reason": null + }, + "variance": { + "status": "measured", + "value": 1e-10, + "units": "J^2", + "normalization": "total", + "reason": null + }, + "magnetization_z": { + "status": "measured", + "value": 0.24931, + "units": "dimensionless", + "normalization": "per-site", + "reason": null + }, + "entanglement_entropy": { + "status": "measured", + "value": [ + 0.132, + 0.284, + 0.315, + 0.284, + 0.132 + ], + "units": "dimensionless", + "normalization": "bond", + "reason": null + }, + "correlator_sigma_x": { + "status": "measured", + "value": [ + 1.0, + 0.731, + 0.548, + 0.402 + ], + "units": "dimensionless", + "normalization": "distance", + "reason": null + } + }, + "convergence": [ + { + "step": 1, + "metrics": { + "energy": -8.749170987567908 + } + }, + { + "step": 2, + "metrics": { + "energy": -8.749170997567909, + "canonical_residual": 1e-10, + "symmetry_residual": 1e-12 + } + } + ], + "warnings": [], + "known_limitations": [ + "Only the finite spin-1/2 TFIM open-chain route is promoted.", + "The fixture does not count as a generated challenge problem." + ] +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-repeat-result.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-repeat-result.json new file mode 100644 index 000000000..387875def --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-repeat-result.json @@ -0,0 +1,140 @@ +{ + "schema_version": "tn-agent.backend-result.v1", + "request_digest": "sha256:247d02d04c82677698405851fe340c440a23b01620870591de0b785a60847a5d", + "binding": { + "schema_version": "tn-agent.backend-binding.v1", + "capability_id": "tenpy.finite_1d.dmrg", + "adapter_id": "tenpy.v1", + "backend_id": "tenpy", + "request_schema": "tn-agent.tenpy.finite-tfim-dmrg.v1", + "result_schema": "tn-agent.backend-result.v1" + }, + "backend": { + "schema_version": "tn-agent.backend-identity.v1", + "backend_id": "tenpy", + "backend_version": "1.1.0", + "adapter_id": "tenpy.v1", + "adapter_version": "1.0.0" + }, + "environment": { + "schema_version": "tn-agent.environment-identity.v1", + "environment_digest": "sha256:33f97f66d496d52d7a68f6a05160f07f90b1d928cea8481f892db7dac132a756", + "runtime_id": "python", + "runtime_version": "3.12.4", + "dependencies": { + "numpy": "2.0.1", + "tenpy": "1.1.0" + } + }, + "execution": { + "schema_version": "tn-agent.execution-evidence.v1", + "status": "succeeded", + "return_code": 0, + "execution_handle": "fixture:finite:repeat", + "retryable": false, + "stdout_digest": "sha256:91cae5519e57dcb3300e0e8934d67448f2ffdc7f45e330bb0b270f4c60ebab8a", + "stderr_digest": "sha256:153ad8c471de56801462ca68e862cd9eb1e40a04cfbf7db67335f0c679c164ec", + "stdout_truncated": false, + "stderr_truncated": false + }, + "observables": { + "correlator_sigma_x": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "correlator_sigma_x", + "status": "measured", + "value": [ + 1.0, + 0.731, + 0.548, + 0.402 + ], + "units": "dimensionless", + "normalization": "distance", + "reason": null + }, + "energy": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "energy", + "status": "measured", + "value": -8.749170997567909, + "units": "J", + "normalization": "total", + "reason": null + }, + "entanglement_entropy": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "entanglement_entropy", + "status": "measured", + "value": [ + 0.132, + 0.284, + 0.315, + 0.284, + 0.132 + ], + "units": "dimensionless", + "normalization": "bond", + "reason": null + }, + "magnetization_z": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "magnetization_z", + "status": "measured", + "value": 0.24931, + "units": "dimensionless", + "normalization": "per-site", + "reason": null + }, + "variance": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "variance", + "status": "measured", + "value": 1e-10, + "units": "J^2", + "normalization": "total", + "reason": null + } + }, + "convergence": [ + { + "schema_version": "tn-agent.convergence-point.v1", + "step": 1, + "metrics": { + "energy": -8.749170987567908 + } + }, + { + "schema_version": "tn-agent.convergence-point.v1", + "step": 2, + "metrics": { + "energy": -8.749170997567909, + "canonical_residual": 1e-10, + "symmetry_residual": 1e-12 + } + } + ], + "diagnostics": [], + "provenance": { + "schema_version": "tn-agent.provenance-evidence.v1", + "plan_id": "sha256:56217e253c4f5e13e66b9939c8ff5f56d26ee77023869288be54c2fe2c0ae307", + "capability_id": "tenpy.finite_1d.dmrg", + "adapter_id": "tenpy.v1", + "raw_result_relative": "backend-repeat-raw-result.json" + }, + "artifacts": [ + { + "schema_version": "tn-agent.result-artifact.v1", + "relative_path": "backend-repeat-raw-result.json", + "digest": "sha256:339d2baf5209c7ce0c50ddb84ee0efabaa000ddad1003725235025acdfb07778", + "media_type": "application/json", + "size_bytes": 2188 + } + ], + "warnings": [], + "known_limitations": [ + "Only the finite spin-1/2 TFIM open-chain route is promoted.", + "The fixture does not count as a generated challenge problem." + ], + "backend_limited_fields": [], + "result_digest": "sha256:4d17aa5b2ae63db1bf04ea42e28be46f0d5a9761d9a1e60f29f297c8ec13bbce" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-repeat-stderr.log b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-repeat-stderr.log new file mode 100644 index 000000000..80610cfd5 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-repeat-stderr.log @@ -0,0 +1 @@ +valid-finite repeat log diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-repeat-stdout.log b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-repeat-stdout.log new file mode 100644 index 000000000..ff5922367 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-repeat-stdout.log @@ -0,0 +1 @@ +valid-finite repeat ok diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-request.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-request.json new file mode 100644 index 000000000..e1a316886 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-request.json @@ -0,0 +1,69 @@ +{ + "schema_version": "tn-agent.tenpy.finite-tfim-dmrg.v1", + "capability_id": "tenpy.finite_1d.dmrg", + "plan_id": "sha256:56217e253c4f5e13e66b9939c8ff5f56d26ee77023869288be54c2fe2c0ae307", + "model_family": "tfim", + "representation": "operator_sum", + "operators": [ + "sigma_x_i*sigma_x_i+1", + "sigma_z_i" + ], + "boundary": "open", + "local_dimension": 2, + "unit_cell": 1, + "algorithm": "dmrg", + "variant": "two_site", + "active_sites": 2, + "initial_state": "all_z_plus", + "allow_bond_growth": true, + "target_precision": 1e-08, + "seed": 1234, + "numerics": { + "max_bond_dim": 32, + "chi_schedule": [ + 32 + ], + "requested_cutoff": 1e-10, + "effective_svd_min": 1e-10, + "engine_cutoff": 1e-10, + "max_sweeps": 5, + "min_sweeps": 0, + "entropy_tolerance": null, + "mixer": true, + "lanczos_maxiter": 8, + "lanczos_n_max": 8, + "checkpoint_every": 2, + "diagonal_gauge_frequency": 0, + "check_overlap": false + }, + "finite_entanglement_fit": { + "enabled": false, + "min_chi": 32, + "max_chi": 128, + "entropy_statistic": "center" + }, + "transfer_matrix": { + "compute": false, + "num_eigs": 4 + }, + "acceptance": { + "require_all_validators": false, + "energy_drift_max": 1e-06, + "variance_max": 0.0, + "canonical_residual_max": 1e-08, + "symmetry_residual_max": 0.0, + "reproducibility_max": 0.0 + }, + "requested_observables": [ + "correlator_sigma_x", + "energy", + "entanglement_entropy", + "magnetization_z", + "variance" + ], + "length": 8, + "coupling_j": 1.0, + "transverse_field_g": 0.8, + "conserve": null, + "target_total_sz": null +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-result.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-result.json new file mode 100644 index 000000000..e1543a208 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-result.json @@ -0,0 +1,140 @@ +{ + "schema_version": "tn-agent.backend-result.v1", + "request_digest": "sha256:247d02d04c82677698405851fe340c440a23b01620870591de0b785a60847a5d", + "binding": { + "schema_version": "tn-agent.backend-binding.v1", + "capability_id": "tenpy.finite_1d.dmrg", + "adapter_id": "tenpy.v1", + "backend_id": "tenpy", + "request_schema": "tn-agent.tenpy.finite-tfim-dmrg.v1", + "result_schema": "tn-agent.backend-result.v1" + }, + "backend": { + "schema_version": "tn-agent.backend-identity.v1", + "backend_id": "tenpy", + "backend_version": "1.1.0", + "adapter_id": "tenpy.v1", + "adapter_version": "1.0.0" + }, + "environment": { + "schema_version": "tn-agent.environment-identity.v1", + "environment_digest": "sha256:33f97f66d496d52d7a68f6a05160f07f90b1d928cea8481f892db7dac132a756", + "runtime_id": "python", + "runtime_version": "3.12.4", + "dependencies": { + "numpy": "2.0.1", + "tenpy": "1.1.0" + } + }, + "execution": { + "schema_version": "tn-agent.execution-evidence.v1", + "status": "succeeded", + "return_code": 0, + "execution_handle": "fixture:finite:primary", + "retryable": false, + "stdout_digest": "sha256:78158f80ea32f0956900cb20ee54117b4c2731b15f2cbc1c44f736343ead468c", + "stderr_digest": "sha256:78bbf8e1c6ab4d84f8d34341d54f7c86eb30bb9ca6a3aeafd8d4d9c5e02ee61f", + "stdout_truncated": false, + "stderr_truncated": false + }, + "observables": { + "correlator_sigma_x": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "correlator_sigma_x", + "status": "measured", + "value": [ + 1.0, + 0.731, + 0.548, + 0.402 + ], + "units": "dimensionless", + "normalization": "distance", + "reason": null + }, + "energy": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "energy", + "status": "measured", + "value": -8.749171017567908, + "units": "J", + "normalization": "total", + "reason": null + }, + "entanglement_entropy": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "entanglement_entropy", + "status": "measured", + "value": [ + 0.132, + 0.284, + 0.315, + 0.284, + 0.132 + ], + "units": "dimensionless", + "normalization": "bond", + "reason": null + }, + "magnetization_z": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "magnetization_z", + "status": "measured", + "value": 0.24931, + "units": "dimensionless", + "normalization": "per-site", + "reason": null + }, + "variance": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "variance", + "status": "measured", + "value": 1e-10, + "units": "J^2", + "normalization": "total", + "reason": null + } + }, + "convergence": [ + { + "schema_version": "tn-agent.convergence-point.v1", + "step": 1, + "metrics": { + "energy": -8.749171007567908 + } + }, + { + "schema_version": "tn-agent.convergence-point.v1", + "step": 2, + "metrics": { + "energy": -8.749171017567908, + "canonical_residual": 1e-10, + "symmetry_residual": 1e-12 + } + } + ], + "diagnostics": [], + "provenance": { + "schema_version": "tn-agent.provenance-evidence.v1", + "plan_id": "sha256:56217e253c4f5e13e66b9939c8ff5f56d26ee77023869288be54c2fe2c0ae307", + "capability_id": "tenpy.finite_1d.dmrg", + "adapter_id": "tenpy.v1", + "raw_result_relative": "backend-raw-result.json" + }, + "artifacts": [ + { + "schema_version": "tn-agent.result-artifact.v1", + "relative_path": "backend-raw-result.json", + "digest": "sha256:bfe7ad8e9de5ad80a643f0a663bdff4301a61564146c44d54814bfe14b82b7d5", + "media_type": "application/json", + "size_bytes": 2188 + } + ], + "warnings": [], + "known_limitations": [ + "Only the finite spin-1/2 TFIM open-chain route is promoted.", + "The fixture does not count as a generated challenge problem." + ], + "backend_limited_fields": [], + "result_digest": "sha256:358e797c846d5ba93d36f6dbf26462f04e10048fc73af82aecc2da9139caf81e" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-stderr.log b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-stderr.log new file mode 100644 index 000000000..a68b3c5c0 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-stderr.log @@ -0,0 +1 @@ +valid-finite primary log diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-stdout.log b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-stdout.log new file mode 100644 index 000000000..8a59e10ee --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/backend-stdout.log @@ -0,0 +1 @@ +valid-finite primary ok diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/energy-reference.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/energy-reference.json new file mode 100644 index 000000000..f77a414c7 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/energy-reference.json @@ -0,0 +1,13 @@ +{ + "schema_version": "wangtheophys.tn-energy-reference.v1", + "reference_id": "finite-tfim-l8-open-j1-g0.8-ed", + "capability_id": "tenpy.finite_1d.dmrg", + "physics_digest": "sha256:71a128d84ecddd84a686a13d971d4bf097a320e1f7eccd7115b3319cf9514a71", + "observable": "energy", + "value": -8.749171017567908, + "units": "J", + "normalization": "total", + "method": "exact_diagonalization", + "citation": "Exact diagonalization anchor for the preregistered open L=8 TFIM Hamiltonian with J=1 and g=0.8.", + "result_digest": "sha256:2de6cfa917e8f4cb7819f60df1135dbed44c60c50d0e39d1cdfe015cce0ec819" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/validator-evidence.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/validator-evidence.json new file mode 100644 index 000000000..774c0e688 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/artifacts/validator-evidence.json @@ -0,0 +1,58 @@ +{ + "schema_version": "wangtheophys.tn-validator-evidence.v1", + "request_digest": "sha256:247d02d04c82677698405851fe340c440a23b01620870591de0b785a60847a5d", + "backend_result_digest": "sha256:358e797c846d5ba93d36f6dbf26462f04e10048fc73af82aecc2da9139caf81e", + "repeat_backend_result_digest": "sha256:4d17aa5b2ae63db1bf04ea42e28be46f0d5a9761d9a1e60f29f297c8ec13bbce", + "reference_artifact_digest": "sha256:4555531abc00ffe7d9c7f0f907ab2def5e5a7c05348a8785a10555ff34e9c134", + "results": [ + { + "id": "parse_consistency", + "metric": null, + "value": null, + "source": "gate.primary_and_repeat_bundle_reconstruction" + }, + { + "id": "convergence", + "metric": "energy_drift", + "value": 1.000000082740371e-08, + "source": "gate.primary_raw.convergence.energy_delta" + }, + { + "id": "variance", + "metric": "variance", + "value": 1e-10, + "source": "reported.primary_raw.observables.variance" + }, + { + "id": "canonical_form", + "metric": "canonical_residual", + "value": 1e-10, + "source": "reported.primary_raw.convergence.canonical_residual" + }, + { + "id": "symmetry_check", + "metric": "symmetry_residual", + "value": 1e-12, + "source": "reported.primary_raw.convergence.symmetry_residual" + }, + { + "id": "benchmark_compare", + "metric": "benchmark_delta", + "value": 0.0, + "source": "gate.primary_energy_vs_preregistered_reference" + }, + { + "id": "reproducibility", + "metric": "reproduction_delta", + "value": 1.999999987845058e-08, + "source": "reported.primary_energy_vs_repeat_raw" + }, + { + "id": "artifact_completeness", + "metric": "missing_artifacts", + "value": 0, + "source": "verified.required_artifact_roles" + } + ], + "result_digest": "sha256:386441bf3399901dc004e2d11fc052631722094be001531d54c940d9b0f5d551" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/evidence.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/evidence.json new file mode 100644 index 000000000..8e33afd11 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/evidence.json @@ -0,0 +1,207 @@ +{ + "schema_version": "wangtheophys.tn-evidence.v1", + "experiment_digest": "sha256:fee6d1d1cd7181b1f4c1f28ce379cce2e756ebcfd806717c13260467485e98f8", + "binding": { + "schema_version": "tn-agent.backend-binding.v1", + "capability_id": "tenpy.finite_1d.dmrg", + "adapter_id": "tenpy.v1", + "backend_id": "tenpy", + "request_schema": "tn-agent.tenpy.finite-tfim-dmrg.v1", + "result_schema": "tn-agent.backend-result.v1" + }, + "execution": { + "schema_version": "tn-agent.execution-evidence.v1", + "status": "succeeded", + "return_code": 0, + "execution_handle": "fixture:finite:primary", + "retryable": false, + "stdout_digest": "sha256:78158f80ea32f0956900cb20ee54117b4c2731b15f2cbc1c44f736343ead468c", + "stderr_digest": "sha256:78bbf8e1c6ab4d84f8d34341d54f7c86eb30bb9ca6a3aeafd8d4d9c5e02ee61f", + "stdout_truncated": false, + "stderr_truncated": false + }, + "repeat_execution": { + "schema_version": "tn-agent.execution-evidence.v1", + "status": "succeeded", + "return_code": 0, + "execution_handle": "fixture:finite:repeat", + "retryable": false, + "stdout_digest": "sha256:91cae5519e57dcb3300e0e8934d67448f2ffdc7f45e330bb0b270f4c60ebab8a", + "stderr_digest": "sha256:153ad8c471de56801462ca68e862cd9eb1e40a04cfbf7db67335f0c679c164ec", + "stdout_truncated": false, + "stderr_truncated": false + }, + "artifacts": [ + { + "relative_path": "backend-request.json", + "digest": "sha256:3639a40e34163739de378171c1ad8887749aafa75c397848689a46ca603d83ef", + "size_bytes": 1635, + "media_type": "application/json", + "role": "backend_request" + }, + { + "relative_path": "backend-raw-result.json", + "digest": "sha256:bfe7ad8e9de5ad80a643f0a663bdff4301a61564146c44d54814bfe14b82b7d5", + "size_bytes": 2188, + "media_type": "application/json", + "role": "backend_raw_result" + }, + { + "relative_path": "backend-result.json", + "digest": "sha256:6976445e93f30c4d0132a9bf4d5d3c1c882639a706f3c281432407d8a3e1c586", + "size_bytes": 4159, + "media_type": "application/json", + "role": "backend_result" + }, + { + "relative_path": "backend-repeat-raw-result.json", + "digest": "sha256:339d2baf5209c7ce0c50ddb84ee0efabaa000ddad1003725235025acdfb07778", + "size_bytes": 2188, + "media_type": "application/json", + "role": "backend_repeat_raw_result" + }, + { + "relative_path": "backend-repeat-result.json", + "digest": "sha256:7414dbd6d0478f4614a7ae1fc08a305c49560c0886dc293264fe22e853fc1b29", + "size_bytes": 4172, + "media_type": "application/json", + "role": "backend_repeat_result" + }, + { + "relative_path": "energy-reference.json", + "digest": "sha256:4555531abc00ffe7d9c7f0f907ab2def5e5a7c05348a8785a10555ff34e9c134", + "size_bytes": 598, + "media_type": "application/json", + "role": "energy_reference" + }, + { + "relative_path": "validator-evidence.json", + "digest": "sha256:10e2a4836f40594d3228f3db4feefd9a040a1936b8010da9433f8a036e8df181", + "size_bytes": 1901, + "media_type": "application/json", + "role": "validator_evidence" + }, + { + "relative_path": "backend-stdout.log", + "digest": "sha256:78158f80ea32f0956900cb20ee54117b4c2731b15f2cbc1c44f736343ead468c", + "size_bytes": 24, + "media_type": "text/plain", + "role": "backend_stdout" + }, + { + "relative_path": "backend-stderr.log", + "digest": "sha256:78bbf8e1c6ab4d84f8d34341d54f7c86eb30bb9ca6a3aeafd8d4d9c5e02ee61f", + "size_bytes": 25, + "media_type": "text/plain", + "role": "backend_stderr" + }, + { + "relative_path": "backend-repeat-stdout.log", + "digest": "sha256:91cae5519e57dcb3300e0e8934d67448f2ffdc7f45e330bb0b270f4c60ebab8a", + "size_bytes": 23, + "media_type": "text/plain", + "role": "backend_repeat_stdout" + }, + { + "relative_path": "backend-repeat-stderr.log", + "digest": "sha256:153ad8c471de56801462ca68e862cd9eb1e40a04cfbf7db67335f0c679c164ec", + "size_bytes": 24, + "media_type": "text/plain", + "role": "backend_repeat_stderr" + } + ], + "observables": [ + { + "name": "energy", + "status": "measured", + "evidence_digest": "sha256:6976445e93f30c4d0132a9bf4d5d3c1c882639a706f3c281432407d8a3e1c586" + }, + { + "name": "variance", + "status": "measured", + "evidence_digest": "sha256:6976445e93f30c4d0132a9bf4d5d3c1c882639a706f3c281432407d8a3e1c586" + }, + { + "name": "magnetization_z", + "status": "measured", + "evidence_digest": "sha256:6976445e93f30c4d0132a9bf4d5d3c1c882639a706f3c281432407d8a3e1c586" + }, + { + "name": "entanglement_entropy", + "status": "measured", + "evidence_digest": "sha256:6976445e93f30c4d0132a9bf4d5d3c1c882639a706f3c281432407d8a3e1c586" + }, + { + "name": "correlator_sigma_x", + "status": "measured", + "evidence_digest": "sha256:6976445e93f30c4d0132a9bf4d5d3c1c882639a706f3c281432407d8a3e1c586" + } + ], + "validator_results": [ + { + "id": "parse_consistency", + "status": "pass", + "reason_code": "VALIDATOR_PASS", + "metric_value": null, + "evidence_digest": "sha256:10e2a4836f40594d3228f3db4feefd9a040a1936b8010da9433f8a036e8df181" + }, + { + "id": "convergence", + "status": "pass", + "reason_code": "VALIDATOR_PASS", + "metric_value": 1.000000082740371e-08, + "evidence_digest": "sha256:10e2a4836f40594d3228f3db4feefd9a040a1936b8010da9433f8a036e8df181" + }, + { + "id": "variance", + "status": "reported_only", + "reason_code": "REPORTED_ONLY", + "metric_value": 1e-10, + "evidence_digest": "sha256:10e2a4836f40594d3228f3db4feefd9a040a1936b8010da9433f8a036e8df181" + }, + { + "id": "canonical_form", + "status": "reported_only", + "reason_code": "REPORTED_ONLY", + "metric_value": 1e-10, + "evidence_digest": "sha256:10e2a4836f40594d3228f3db4feefd9a040a1936b8010da9433f8a036e8df181" + }, + { + "id": "symmetry_check", + "status": "reported_only", + "reason_code": "REPORTED_ONLY", + "metric_value": 1e-12, + "evidence_digest": "sha256:10e2a4836f40594d3228f3db4feefd9a040a1936b8010da9433f8a036e8df181" + }, + { + "id": "benchmark_compare", + "status": "pass", + "reason_code": "VALIDATOR_PASS", + "metric_value": 0.0, + "evidence_digest": "sha256:10e2a4836f40594d3228f3db4feefd9a040a1936b8010da9433f8a036e8df181" + }, + { + "id": "reproducibility", + "status": "reported_only", + "reason_code": "REPORTED_ONLY", + "metric_value": 1.999999987845058e-08, + "evidence_digest": "sha256:10e2a4836f40594d3228f3db4feefd9a040a1936b8010da9433f8a036e8df181" + }, + { + "id": "artifact_completeness", + "status": "pass", + "reason_code": "VALIDATOR_PASS", + "metric_value": 0, + "evidence_digest": "sha256:10e2a4836f40594d3228f3db4feefd9a040a1936b8010da9433f8a036e8df181" + } + ], + "provenance": { + "plan_id": "sha256:56217e253c4f5e13e66b9939c8ff5f56d26ee77023869288be54c2fe2c0ae307", + "request_digest": "sha256:247d02d04c82677698405851fe340c440a23b01620870591de0b785a60847a5d", + "backend_result_digest": "sha256:6976445e93f30c4d0132a9bf4d5d3c1c882639a706f3c281432407d8a3e1c586", + "repeat_backend_result_digest": "sha256:7414dbd6d0478f4614a7ae1fc08a305c49560c0886dc293264fe22e853fc1b29", + "generated_by": "fixtures/regenerate.py", + "generated_at": "2026-07-29T00:01:00Z" + }, + "result_digest": "sha256:f7f261e1ae0222ea64ceaa6a295dfea9cda283d14eb86c78deab750c0d819b76" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/experiment.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/experiment.json new file mode 100644 index 000000000..f9173e8f7 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-finite/experiment.json @@ -0,0 +1,193 @@ +{ + "schema_version": "wangtheophys.tn-experiment.v1", + "problem": { + "problem_id": "fixture-finite-tfim-dmrg", + "title": "Finite TFIM DMRG contract fixture", + "research_question": "Can the promoted finite TFIM route produce evidence that satisfies its preregistered numerical and provenance gates?", + "novelty_claim": "This is a contract fixture, not a claim of a new scientific result.", + "status": "test_fixture", + "task_family": "ground_state_1d_finite" + }, + "physics": { + "model": { + "representation": "operator_sum", + "family": "tfim", + "operators": [ + "sigma_x_i*sigma_x_i+1", + "sigma_z_i" + ], + "couplings": { + "J": 1.0, + "g": 0.8 + }, + "onsite_terms": [], + "neighbor_range": 1 + }, + "lattice": { + "type": "chain", + "length": 8, + "boundary": "open", + "local_dim": 2, + "unit_cell": 1 + }, + "symmetry": { + "U1_Sz": false, + "parity": false, + "translation": false, + "target_sector": null + }, + "ansatz": { + "family": "mps", + "algorithm": "dmrg", + "variant": "two_site", + "initial_state": "all_z_plus", + "allow_bond_growth": true, + "target_precision": 1e-08 + } + }, + "capability": { + "capability_id": "tenpy.finite_1d.dmrg", + "maturity": "stable", + "known_limitations": [ + "Only the finite spin-1/2 TFIM open-chain route is promoted.", + "The fixture does not count as a generated challenge problem." + ] + }, + "backend_binding": { + "schema_version": "tn-agent.backend-binding.v1", + "capability_id": "tenpy.finite_1d.dmrg", + "adapter_id": "tenpy.v1", + "backend_id": "tenpy", + "request_schema": "tn-agent.tenpy.finite-tfim-dmrg.v1", + "result_schema": "tn-agent.backend-result.v1" + }, + "numerics": { + "max_bond_dim": 32, + "cutoff": 1e-10, + "max_sweeps": 5, + "mixer": true, + "lanczos_maxiter": 8, + "checkpoint_every": 2, + "finite_entanglement_fit": { + "enabled": false, + "min_chi": 32, + "max_chi": 128 + }, + "transfer_matrix": { + "compute": false, + "num_eigs": 4 + }, + "seed": 1234 + }, + "observables": [ + "energy", + "variance", + "magnetization_z", + "entanglement_entropy", + "correlator_sigma_x" + ], + "validators": [ + { + "id": "parse_consistency", + "policy": "required_pass", + "metric": null, + "operator": null, + "threshold": null + }, + { + "id": "convergence", + "policy": "required_pass", + "metric": "energy_drift", + "operator": "max", + "threshold": 1e-06 + }, + { + "id": "variance", + "policy": "reported_only", + "metric": "variance", + "operator": null, + "threshold": null + }, + { + "id": "canonical_form", + "policy": "reported_only", + "metric": "canonical_residual", + "operator": null, + "threshold": null + }, + { + "id": "symmetry_check", + "policy": "reported_only", + "metric": "symmetry_residual", + "operator": null, + "threshold": null + }, + { + "id": "benchmark_compare", + "policy": "required_pass", + "metric": "benchmark_delta", + "operator": "max", + "threshold": 0.001 + }, + { + "id": "reproducibility", + "policy": "reported_only", + "metric": "reproduction_delta", + "operator": null, + "threshold": null + }, + { + "id": "artifact_completeness", + "policy": "required_pass", + "metric": "missing_artifacts", + "operator": "equals", + "threshold": 0 + } + ], + "acceptance": { + "mode": "all_required", + "required_validator_ids": [ + "parse_consistency", + "convergence", + "benchmark_compare", + "artifact_completeness" + ], + "allowed_backend_limited_ids": [], + "require_execution_success": true, + "reported_only_validator_ids": [ + "variance", + "canonical_form", + "symmetry_check", + "reproducibility" + ] + }, + "provenance": { + "created_by": "WangTheoPhys", + "created_at": "2026-07-29T00:00:00Z", + "sources": [ + { + "kind": "repository_skill", + "uri": "skills/method-mps/SKILL.md", + "sha256": "sha256:297e1df7bb54e403edf86b4e5b683a6774d0db78673cd0fc76dc116332dbf088" + }, + { + "kind": "repository_skill", + "uri": "skills/using-tenpy/SKILL.md", + "sha256": "sha256:622b88619105cc0643bfde6d295377dd0c90ceee297119029362e47e80c14f5d" + } + ], + "generation_log_uri": "tracks/agent-kb/solutions/WangTheoPhys/README.md", + "human_gatekeeper_role": "Ratify the Hamiltonian, geometry, symmetry sector, target observable, and preregistered acceptance gate before execution." + }, + "reference": { + "observable": "energy", + "value": -8.749171017567908, + "units": "J", + "normalization": "total", + "source": { + "kind": "registered_artifact", + "uri": "energy-reference.json", + "sha256": "sha256:4555531abc00ffe7d9c7f0f907ab2def5e5a7c05348a8785a10555ff34e9c134" + } + } +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-raw-result.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-raw-result.json new file mode 100644 index 000000000..cd0a25052 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-raw-result.json @@ -0,0 +1,90 @@ +{ + "schema_version": "tn-agent.tenpy.raw-result.v1", + "request_digest": "sha256:e8a0eed66378a47b078a99227b308775c34312b561efbb0988fd6e8beff4c064", + "plan_id": "sha256:583e130a0daa654eeecfe215f84196d37cb1792480d811dfc42f11b1cb25df60", + "capability_id": "tenpy.infinite_1d.vumps", + "backend_id": "tenpy", + "backend_version": "1.1.0", + "adapter_id": "tenpy.v1", + "adapter_version": "1.0.0", + "status": "succeeded", + "environment": { + "schema_version": "tn-agent.tenpy.environment.v1", + "environment_digest": "sha256:33f97f66d496d52d7a68f6a05160f07f90b1d928cea8481f892db7dac132a756", + "runtime_id": "python", + "runtime_version": "3.12.4", + "dependencies": { + "numpy": "2.0.1", + "tenpy": "1.1.0" + } + }, + "observables": { + "energy": { + "status": "measured", + "value": -0.4431, + "units": "J", + "normalization": "per-site", + "reason": null + }, + "variance": { + "status": "backend_limited", + "value": null, + "units": "J^2", + "normalization": "per-site", + "reason": "TeNPy 1.1.0 does not expose UniformMPS energy variance" + }, + "entanglement_entropy": { + "status": "measured", + "value": [ + 0.612, + 0.731 + ], + "units": "dimensionless", + "normalization": "bond", + "reason": null + }, + "transfer_spectrum": { + "status": "measured", + "value": [ + 1.0, + 0.721, + 0.519, + 0.374 + ], + "units": "dimensionless", + "normalization": "spectrum", + "reason": null + }, + "central_charge_fit": { + "status": "derived", + "value": 0.501, + "units": "dimensionless", + "normalization": "fit", + "reason": null + } + }, + "convergence": [ + { + "step": 1, + "metrics": { + "energy": -0.443 + } + }, + { + "step": 2, + "metrics": { + "energy": -0.4431, + "canonical_residual": 0.001, + "symmetry_residual": 1e-10 + } + } + ], + "warnings": [ + "variance is backend-limited on the infinite route" + ], + "known_limitations": [ + "TeNPy UniformMPS is experimental.", + "Energy variance is backend-limited on this route.", + "The fixture does not count as a generated challenge problem." + ] +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-repeat-raw-result.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-repeat-raw-result.json new file mode 100644 index 000000000..011407f0f --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-repeat-raw-result.json @@ -0,0 +1,90 @@ +{ + "schema_version": "tn-agent.tenpy.raw-result.v1", + "request_digest": "sha256:e8a0eed66378a47b078a99227b308775c34312b561efbb0988fd6e8beff4c064", + "plan_id": "sha256:583e130a0daa654eeecfe215f84196d37cb1792480d811dfc42f11b1cb25df60", + "capability_id": "tenpy.infinite_1d.vumps", + "backend_id": "tenpy", + "backend_version": "1.1.0", + "adapter_id": "tenpy.v1", + "adapter_version": "1.0.0", + "status": "succeeded", + "environment": { + "schema_version": "tn-agent.tenpy.environment.v1", + "environment_digest": "sha256:33f97f66d496d52d7a68f6a05160f07f90b1d928cea8481f892db7dac132a756", + "runtime_id": "python", + "runtime_version": "3.12.4", + "dependencies": { + "numpy": "2.0.1", + "tenpy": "1.1.0" + } + }, + "observables": { + "energy": { + "status": "measured", + "value": -0.44311, + "units": "J", + "normalization": "per-site", + "reason": null + }, + "variance": { + "status": "backend_limited", + "value": null, + "units": "J^2", + "normalization": "per-site", + "reason": "TeNPy 1.1.0 does not expose UniformMPS energy variance" + }, + "entanglement_entropy": { + "status": "measured", + "value": [ + 0.612, + 0.731 + ], + "units": "dimensionless", + "normalization": "bond", + "reason": null + }, + "transfer_spectrum": { + "status": "measured", + "value": [ + 1.0, + 0.721, + 0.519, + 0.374 + ], + "units": "dimensionless", + "normalization": "spectrum", + "reason": null + }, + "central_charge_fit": { + "status": "derived", + "value": 0.501, + "units": "dimensionless", + "normalization": "fit", + "reason": null + } + }, + "convergence": [ + { + "step": 1, + "metrics": { + "energy": -0.44301 + } + }, + { + "step": 2, + "metrics": { + "energy": -0.44311, + "canonical_residual": 0.001, + "symmetry_residual": 1e-10 + } + } + ], + "warnings": [ + "variance is backend-limited on the infinite route" + ], + "known_limitations": [ + "TeNPy UniformMPS is experimental.", + "Energy variance is backend-limited on this route.", + "The fixture does not count as a generated challenge problem." + ] +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-repeat-result.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-repeat-result.json new file mode 100644 index 000000000..8fe16f740 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-repeat-result.json @@ -0,0 +1,142 @@ +{ + "schema_version": "tn-agent.backend-result.v1", + "request_digest": "sha256:e8a0eed66378a47b078a99227b308775c34312b561efbb0988fd6e8beff4c064", + "binding": { + "schema_version": "tn-agent.backend-binding.v1", + "capability_id": "tenpy.infinite_1d.vumps", + "adapter_id": "tenpy.v1", + "backend_id": "tenpy", + "request_schema": "tn-agent.tenpy.infinite-xxz-vumps.v1", + "result_schema": "tn-agent.backend-result.v1" + }, + "backend": { + "schema_version": "tn-agent.backend-identity.v1", + "backend_id": "tenpy", + "backend_version": "1.1.0", + "adapter_id": "tenpy.v1", + "adapter_version": "1.0.0" + }, + "environment": { + "schema_version": "tn-agent.environment-identity.v1", + "environment_digest": "sha256:33f97f66d496d52d7a68f6a05160f07f90b1d928cea8481f892db7dac132a756", + "runtime_id": "python", + "runtime_version": "3.12.4", + "dependencies": { + "numpy": "2.0.1", + "tenpy": "1.1.0" + } + }, + "execution": { + "schema_version": "tn-agent.execution-evidence.v1", + "status": "succeeded", + "return_code": 0, + "execution_handle": "fixture:infinite:repeat", + "retryable": false, + "stdout_digest": "sha256:f257797cb41fe5262fa60cee3d2f7d559fb7a28d1cdd6938648567299883349a", + "stderr_digest": "sha256:d034e30db09af58a309663aa53bc609b60f4f06d4f5fa805756f9df16f327051", + "stdout_truncated": false, + "stderr_truncated": false + }, + "observables": { + "central_charge_fit": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "central_charge_fit", + "status": "derived", + "value": 0.501, + "units": "dimensionless", + "normalization": "fit", + "reason": null + }, + "energy": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "energy", + "status": "measured", + "value": -0.44311, + "units": "J", + "normalization": "per-site", + "reason": null + }, + "entanglement_entropy": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "entanglement_entropy", + "status": "measured", + "value": [ + 0.612, + 0.731 + ], + "units": "dimensionless", + "normalization": "bond", + "reason": null + }, + "transfer_spectrum": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "transfer_spectrum", + "status": "measured", + "value": [ + 1.0, + 0.721, + 0.519, + 0.374 + ], + "units": "dimensionless", + "normalization": "spectrum", + "reason": null + }, + "variance": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "variance", + "status": "backend_limited", + "value": null, + "units": "J^2", + "normalization": "per-site", + "reason": "TeNPy 1.1.0 does not expose UniformMPS energy variance" + } + }, + "convergence": [ + { + "schema_version": "tn-agent.convergence-point.v1", + "step": 1, + "metrics": { + "energy": -0.44301 + } + }, + { + "schema_version": "tn-agent.convergence-point.v1", + "step": 2, + "metrics": { + "energy": -0.44311, + "canonical_residual": 0.001, + "symmetry_residual": 1e-10 + } + } + ], + "diagnostics": [], + "provenance": { + "schema_version": "tn-agent.provenance-evidence.v1", + "plan_id": "sha256:583e130a0daa654eeecfe215f84196d37cb1792480d811dfc42f11b1cb25df60", + "capability_id": "tenpy.infinite_1d.vumps", + "adapter_id": "tenpy.v1", + "raw_result_relative": "backend-repeat-raw-result.json" + }, + "artifacts": [ + { + "schema_version": "tn-agent.result-artifact.v1", + "relative_path": "backend-repeat-raw-result.json", + "digest": "sha256:caa3d44cb7272c0e44118bd413bc68fce1551385fa17c370c00d05761fcecb15", + "media_type": "application/json", + "size_bytes": 2264 + } + ], + "warnings": [ + "variance is backend-limited on the infinite route" + ], + "known_limitations": [ + "TeNPy UniformMPS is experimental.", + "Energy variance is backend-limited on this route.", + "The fixture does not count as a generated challenge problem." + ], + "backend_limited_fields": [ + "variance" + ], + "result_digest": "sha256:f166be6c9aa8fad1aa62e260b04e2e54328756f2fc89b653d106acebae2d6c50" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-repeat-stderr.log b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-repeat-stderr.log new file mode 100644 index 000000000..bcc942bed --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-repeat-stderr.log @@ -0,0 +1 @@ +valid-infinite repeat log diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-repeat-stdout.log b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-repeat-stdout.log new file mode 100644 index 000000000..af506f70e --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-repeat-stdout.log @@ -0,0 +1 @@ +valid-infinite repeat ok diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-request.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-request.json new file mode 100644 index 000000000..16b132be5 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-request.json @@ -0,0 +1,72 @@ +{ + "schema_version": "tn-agent.tenpy.infinite-xxz-vumps.v1", + "capability_id": "tenpy.infinite_1d.vumps", + "plan_id": "sha256:583e130a0daa654eeecfe215f84196d37cb1792480d811dfc42f11b1cb25df60", + "model_family": "xxz", + "representation": "operator_sum", + "operators": [ + "Sx_i*Sx_i+1", + "Sy_i*Sy_i+1", + "Sz_i", + "Sz_i*Sz_i+1" + ], + "boundary": "infinite", + "local_dimension": 2, + "unit_cell": 2, + "algorithm": "vumps", + "variant": "two_site", + "active_sites": 2, + "initial_state": "neel", + "allow_bond_growth": true, + "target_precision": 1e-06, + "seed": 2468, + "numerics": { + "max_bond_dim": 32, + "chi_schedule": [ + 16, + 32 + ], + "requested_cutoff": 1e-10, + "effective_svd_min": null, + "engine_cutoff": 0.0, + "max_sweeps": 4, + "min_sweeps": 2, + "entropy_tolerance": 1e-05, + "mixer": false, + "lanczos_maxiter": 8, + "lanczos_n_max": 32, + "checkpoint_every": 1, + "diagonal_gauge_frequency": 0, + "check_overlap": false + }, + "finite_entanglement_fit": { + "enabled": true, + "min_chi": 16, + "max_chi": 32, + "entropy_statistic": "center" + }, + "transfer_matrix": { + "compute": true, + "num_eigs": 4 + }, + "acceptance": { + "require_all_validators": false, + "energy_drift_max": 0.0002, + "variance_max": 0.0, + "canonical_residual_max": 1e-06, + "symmetry_residual_max": 0.0, + "reproducibility_max": 0.0 + }, + "requested_observables": [ + "central_charge_fit", + "energy", + "entanglement_entropy", + "transfer_spectrum", + "variance" + ], + "coupling_jxy": 1.0, + "anisotropy_delta": 1.0, + "field_h": 0.0, + "conserve": "Sz", + "target_total_sz": 0 +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-result.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-result.json new file mode 100644 index 000000000..28afa2111 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-result.json @@ -0,0 +1,142 @@ +{ + "schema_version": "tn-agent.backend-result.v1", + "request_digest": "sha256:e8a0eed66378a47b078a99227b308775c34312b561efbb0988fd6e8beff4c064", + "binding": { + "schema_version": "tn-agent.backend-binding.v1", + "capability_id": "tenpy.infinite_1d.vumps", + "adapter_id": "tenpy.v1", + "backend_id": "tenpy", + "request_schema": "tn-agent.tenpy.infinite-xxz-vumps.v1", + "result_schema": "tn-agent.backend-result.v1" + }, + "backend": { + "schema_version": "tn-agent.backend-identity.v1", + "backend_id": "tenpy", + "backend_version": "1.1.0", + "adapter_id": "tenpy.v1", + "adapter_version": "1.0.0" + }, + "environment": { + "schema_version": "tn-agent.environment-identity.v1", + "environment_digest": "sha256:33f97f66d496d52d7a68f6a05160f07f90b1d928cea8481f892db7dac132a756", + "runtime_id": "python", + "runtime_version": "3.12.4", + "dependencies": { + "numpy": "2.0.1", + "tenpy": "1.1.0" + } + }, + "execution": { + "schema_version": "tn-agent.execution-evidence.v1", + "status": "succeeded", + "return_code": 0, + "execution_handle": "fixture:infinite:primary", + "retryable": false, + "stdout_digest": "sha256:ba5226f2581a284e5171d6f235c08eeb537709c4dbcb0c3a90b4530a7dd633a1", + "stderr_digest": "sha256:76dcbfe7624517874bbf3765410e2f588877eb80ad42e92a823d2c21d3c94add", + "stdout_truncated": false, + "stderr_truncated": false + }, + "observables": { + "central_charge_fit": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "central_charge_fit", + "status": "derived", + "value": 0.501, + "units": "dimensionless", + "normalization": "fit", + "reason": null + }, + "energy": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "energy", + "status": "measured", + "value": -0.4431, + "units": "J", + "normalization": "per-site", + "reason": null + }, + "entanglement_entropy": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "entanglement_entropy", + "status": "measured", + "value": [ + 0.612, + 0.731 + ], + "units": "dimensionless", + "normalization": "bond", + "reason": null + }, + "transfer_spectrum": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "transfer_spectrum", + "status": "measured", + "value": [ + 1.0, + 0.721, + 0.519, + 0.374 + ], + "units": "dimensionless", + "normalization": "spectrum", + "reason": null + }, + "variance": { + "schema_version": "tn-agent.observable-evidence.v1", + "name": "variance", + "status": "backend_limited", + "value": null, + "units": "J^2", + "normalization": "per-site", + "reason": "TeNPy 1.1.0 does not expose UniformMPS energy variance" + } + }, + "convergence": [ + { + "schema_version": "tn-agent.convergence-point.v1", + "step": 1, + "metrics": { + "energy": -0.443 + } + }, + { + "schema_version": "tn-agent.convergence-point.v1", + "step": 2, + "metrics": { + "energy": -0.4431, + "canonical_residual": 0.001, + "symmetry_residual": 1e-10 + } + } + ], + "diagnostics": [], + "provenance": { + "schema_version": "tn-agent.provenance-evidence.v1", + "plan_id": "sha256:583e130a0daa654eeecfe215f84196d37cb1792480d811dfc42f11b1cb25df60", + "capability_id": "tenpy.infinite_1d.vumps", + "adapter_id": "tenpy.v1", + "raw_result_relative": "backend-raw-result.json" + }, + "artifacts": [ + { + "schema_version": "tn-agent.result-artifact.v1", + "relative_path": "backend-raw-result.json", + "digest": "sha256:ef892045230c418b022ab3e43b33e2b4046ed9790aca62bb3161675e2ecbc38c", + "media_type": "application/json", + "size_bytes": 2260 + } + ], + "warnings": [ + "variance is backend-limited on the infinite route" + ], + "known_limitations": [ + "TeNPy UniformMPS is experimental.", + "Energy variance is backend-limited on this route.", + "The fixture does not count as a generated challenge problem." + ], + "backend_limited_fields": [ + "variance" + ], + "result_digest": "sha256:7f710c61dc85c250b0d77d5a98ebb1a51194ccbd18cb80df89be885392a9437e" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-stderr.log b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-stderr.log new file mode 100644 index 000000000..f63e1ffcb --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-stderr.log @@ -0,0 +1 @@ +valid-infinite primary log diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-stdout.log b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-stdout.log new file mode 100644 index 000000000..0f771623f --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/backend-stdout.log @@ -0,0 +1 @@ +valid-infinite primary ok diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/energy-reference.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/energy-reference.json new file mode 100644 index 000000000..2228899b9 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/energy-reference.json @@ -0,0 +1,13 @@ +{ + "schema_version": "wangtheophys.tn-energy-reference.v1", + "reference_id": "infinite-heisenberg-delta1-bethe-ansatz", + "capability_id": "tenpy.infinite_1d.vumps", + "physics_digest": "sha256:877bcf8045fe0aa5dcebbb6aa310eeab144da3a2aea21ddd8bdc591bcab649bf", + "observable": "energy", + "value": -0.4431471805599453, + "units": "J", + "normalization": "per-site", + "method": "analytic_bethe_ansatz", + "citation": "Thermodynamic-limit spin-1/2 antiferromagnetic Heisenberg-chain ground-state energy per site, 1/4-ln(2).", + "result_digest": "sha256:a509283a89f7e508f3f8755f7bb34582bcacff7ec53362541e56e30f4ec59ce2" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/validator-evidence.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/validator-evidence.json new file mode 100644 index 000000000..0e50386cd --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/artifacts/validator-evidence.json @@ -0,0 +1,58 @@ +{ + "schema_version": "wangtheophys.tn-validator-evidence.v1", + "request_digest": "sha256:e8a0eed66378a47b078a99227b308775c34312b561efbb0988fd6e8beff4c064", + "backend_result_digest": "sha256:7f710c61dc85c250b0d77d5a98ebb1a51194ccbd18cb80df89be885392a9437e", + "repeat_backend_result_digest": "sha256:f166be6c9aa8fad1aa62e260b04e2e54328756f2fc89b653d106acebae2d6c50", + "reference_artifact_digest": "sha256:6cacc4617490466b747b041270e6316ad90732c342441949bac118fc84652ef3", + "results": [ + { + "id": "parse_consistency", + "metric": null, + "value": null, + "source": "gate.primary_and_repeat_bundle_reconstruction" + }, + { + "id": "convergence", + "metric": "energy_drift", + "value": 9.999999999998899e-05, + "source": "gate.primary_raw.convergence.energy_delta" + }, + { + "id": "variance", + "metric": null, + "value": null, + "source": "reported.primary_raw.observables.variance" + }, + { + "id": "canonical_form", + "metric": "canonical_residual", + "value": 0.001, + "source": "reported.primary_raw.convergence.canonical_residual" + }, + { + "id": "symmetry_check", + "metric": "symmetry_residual", + "value": 1e-10, + "source": "reported.primary_raw.convergence.symmetry_residual" + }, + { + "id": "benchmark_compare", + "metric": "benchmark_delta", + "value": 4.7180559945292355e-05, + "source": "gate.primary_energy_vs_preregistered_reference" + }, + { + "id": "reproducibility", + "metric": "reproduction_delta", + "value": 1.0000000000010001e-05, + "source": "reported.primary_energy_vs_repeat_raw" + }, + { + "id": "artifact_completeness", + "metric": "missing_artifacts", + "value": 0, + "source": "verified.required_artifact_roles" + } + ], + "result_digest": "sha256:06821077dc75b6573e4e85792cce37bffdab08878f7233bc446801277902d9a7" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/evidence.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/evidence.json new file mode 100644 index 000000000..ebffc20ff --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/evidence.json @@ -0,0 +1,207 @@ +{ + "schema_version": "wangtheophys.tn-evidence.v1", + "experiment_digest": "sha256:09a3d34c8886fbf75ae0cebe324e916a9ac2d096fcf5546c22542131dcd38402", + "binding": { + "schema_version": "tn-agent.backend-binding.v1", + "capability_id": "tenpy.infinite_1d.vumps", + "adapter_id": "tenpy.v1", + "backend_id": "tenpy", + "request_schema": "tn-agent.tenpy.infinite-xxz-vumps.v1", + "result_schema": "tn-agent.backend-result.v1" + }, + "execution": { + "schema_version": "tn-agent.execution-evidence.v1", + "status": "succeeded", + "return_code": 0, + "execution_handle": "fixture:infinite:primary", + "retryable": false, + "stdout_digest": "sha256:ba5226f2581a284e5171d6f235c08eeb537709c4dbcb0c3a90b4530a7dd633a1", + "stderr_digest": "sha256:76dcbfe7624517874bbf3765410e2f588877eb80ad42e92a823d2c21d3c94add", + "stdout_truncated": false, + "stderr_truncated": false + }, + "repeat_execution": { + "schema_version": "tn-agent.execution-evidence.v1", + "status": "succeeded", + "return_code": 0, + "execution_handle": "fixture:infinite:repeat", + "retryable": false, + "stdout_digest": "sha256:f257797cb41fe5262fa60cee3d2f7d559fb7a28d1cdd6938648567299883349a", + "stderr_digest": "sha256:d034e30db09af58a309663aa53bc609b60f4f06d4f5fa805756f9df16f327051", + "stdout_truncated": false, + "stderr_truncated": false + }, + "artifacts": [ + { + "relative_path": "backend-request.json", + "digest": "sha256:a5566b0421829140ead78236f3a84e655d44c3e6f8167427fe0f429cd7583d97", + "size_bytes": 1671, + "media_type": "application/json", + "role": "backend_request" + }, + { + "relative_path": "backend-raw-result.json", + "digest": "sha256:ef892045230c418b022ab3e43b33e2b4046ed9790aca62bb3161675e2ecbc38c", + "size_bytes": 2260, + "media_type": "application/json", + "role": "backend_raw_result" + }, + { + "relative_path": "backend-result.json", + "digest": "sha256:25af845170948a00ba3c3a64ed705cb5051f145d979c68fc1fad00f00b220ac7", + "size_bytes": 4258, + "media_type": "application/json", + "role": "backend_result" + }, + { + "relative_path": "backend-repeat-raw-result.json", + "digest": "sha256:caa3d44cb7272c0e44118bd413bc68fce1551385fa17c370c00d05761fcecb15", + "size_bytes": 2264, + "media_type": "application/json", + "role": "backend_repeat_raw_result" + }, + { + "relative_path": "backend-repeat-result.json", + "digest": "sha256:d8afffee85201b2392b8bf87d9f5c8b318d72eafd860142dc06fccbf0903ddad", + "size_bytes": 4275, + "media_type": "application/json", + "role": "backend_repeat_result" + }, + { + "relative_path": "energy-reference.json", + "digest": "sha256:6cacc4617490466b747b041270e6316ad90732c342441949bac118fc84652ef3", + "size_bytes": 622, + "media_type": "application/json", + "role": "energy_reference" + }, + { + "relative_path": "validator-evidence.json", + "digest": "sha256:f63f87e3c0b467056aa4c379d135148db8ab158a8127cd592a3350a96f5ff818", + "size_bytes": 1914, + "media_type": "application/json", + "role": "validator_evidence" + }, + { + "relative_path": "backend-stdout.log", + "digest": "sha256:ba5226f2581a284e5171d6f235c08eeb537709c4dbcb0c3a90b4530a7dd633a1", + "size_bytes": 26, + "media_type": "text/plain", + "role": "backend_stdout" + }, + { + "relative_path": "backend-stderr.log", + "digest": "sha256:76dcbfe7624517874bbf3765410e2f588877eb80ad42e92a823d2c21d3c94add", + "size_bytes": 27, + "media_type": "text/plain", + "role": "backend_stderr" + }, + { + "relative_path": "backend-repeat-stdout.log", + "digest": "sha256:f257797cb41fe5262fa60cee3d2f7d559fb7a28d1cdd6938648567299883349a", + "size_bytes": 25, + "media_type": "text/plain", + "role": "backend_repeat_stdout" + }, + { + "relative_path": "backend-repeat-stderr.log", + "digest": "sha256:d034e30db09af58a309663aa53bc609b60f4f06d4f5fa805756f9df16f327051", + "size_bytes": 26, + "media_type": "text/plain", + "role": "backend_repeat_stderr" + } + ], + "observables": [ + { + "name": "energy", + "status": "measured", + "evidence_digest": "sha256:25af845170948a00ba3c3a64ed705cb5051f145d979c68fc1fad00f00b220ac7" + }, + { + "name": "variance", + "status": "backend_limited", + "evidence_digest": "sha256:25af845170948a00ba3c3a64ed705cb5051f145d979c68fc1fad00f00b220ac7" + }, + { + "name": "entanglement_entropy", + "status": "measured", + "evidence_digest": "sha256:25af845170948a00ba3c3a64ed705cb5051f145d979c68fc1fad00f00b220ac7" + }, + { + "name": "transfer_spectrum", + "status": "measured", + "evidence_digest": "sha256:25af845170948a00ba3c3a64ed705cb5051f145d979c68fc1fad00f00b220ac7" + }, + { + "name": "central_charge_fit", + "status": "derived", + "evidence_digest": "sha256:25af845170948a00ba3c3a64ed705cb5051f145d979c68fc1fad00f00b220ac7" + } + ], + "validator_results": [ + { + "id": "parse_consistency", + "status": "pass", + "reason_code": "VALIDATOR_PASS", + "metric_value": null, + "evidence_digest": "sha256:f63f87e3c0b467056aa4c379d135148db8ab158a8127cd592a3350a96f5ff818" + }, + { + "id": "convergence", + "status": "pass", + "reason_code": "VALIDATOR_PASS", + "metric_value": 9.999999999998899e-05, + "evidence_digest": "sha256:f63f87e3c0b467056aa4c379d135148db8ab158a8127cd592a3350a96f5ff818" + }, + { + "id": "variance", + "status": "backend_limited", + "reason_code": "BACKEND_LIMITED", + "metric_value": null, + "evidence_digest": "sha256:f63f87e3c0b467056aa4c379d135148db8ab158a8127cd592a3350a96f5ff818" + }, + { + "id": "canonical_form", + "status": "reported_only", + "reason_code": "REPORTED_ONLY", + "metric_value": 0.001, + "evidence_digest": "sha256:f63f87e3c0b467056aa4c379d135148db8ab158a8127cd592a3350a96f5ff818" + }, + { + "id": "symmetry_check", + "status": "reported_only", + "reason_code": "REPORTED_ONLY", + "metric_value": 1e-10, + "evidence_digest": "sha256:f63f87e3c0b467056aa4c379d135148db8ab158a8127cd592a3350a96f5ff818" + }, + { + "id": "benchmark_compare", + "status": "pass", + "reason_code": "VALIDATOR_PASS", + "metric_value": 4.7180559945292355e-05, + "evidence_digest": "sha256:f63f87e3c0b467056aa4c379d135148db8ab158a8127cd592a3350a96f5ff818" + }, + { + "id": "reproducibility", + "status": "reported_only", + "reason_code": "REPORTED_ONLY", + "metric_value": 1.0000000000010001e-05, + "evidence_digest": "sha256:f63f87e3c0b467056aa4c379d135148db8ab158a8127cd592a3350a96f5ff818" + }, + { + "id": "artifact_completeness", + "status": "pass", + "reason_code": "VALIDATOR_PASS", + "metric_value": 0, + "evidence_digest": "sha256:f63f87e3c0b467056aa4c379d135148db8ab158a8127cd592a3350a96f5ff818" + } + ], + "provenance": { + "plan_id": "sha256:583e130a0daa654eeecfe215f84196d37cb1792480d811dfc42f11b1cb25df60", + "request_digest": "sha256:e8a0eed66378a47b078a99227b308775c34312b561efbb0988fd6e8beff4c064", + "backend_result_digest": "sha256:25af845170948a00ba3c3a64ed705cb5051f145d979c68fc1fad00f00b220ac7", + "repeat_backend_result_digest": "sha256:d8afffee85201b2392b8bf87d9f5c8b318d72eafd860142dc06fccbf0903ddad", + "generated_by": "fixtures/regenerate.py", + "generated_at": "2026-07-29T00:01:00Z" + }, + "result_digest": "sha256:ed178a5c36c76fdc3c4e94f17a095b6e979bc8cf1067e16efefa49af7b430a01" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/experiment.json b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/experiment.json new file mode 100644 index 000000000..a5f3c7e09 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/fixtures/valid-infinite/experiment.json @@ -0,0 +1,202 @@ +{ + "schema_version": "wangtheophys.tn-experiment.v1", + "problem": { + "problem_id": "fixture-infinite-xxz-vumps", + "title": "Infinite XXZ VUMPS contract fixture", + "research_question": "Can the promoted infinite XXZ route produce evidence that satisfies its preregistered numerical and provenance gates while declaring backend-limited variance?", + "novelty_claim": "This is a contract fixture, not a claim of a new scientific result.", + "status": "test_fixture", + "task_family": "ground_state_1d_infinite" + }, + "physics": { + "model": { + "representation": "operator_sum", + "family": "xxz", + "operators": [ + "Sx_i*Sx_i+1", + "Sy_i*Sy_i+1", + "Sz_i*Sz_i+1", + "Sz_i" + ], + "couplings": { + "Jxy": 1.0, + "Delta": 1.0, + "h": 0.0 + }, + "onsite_terms": [], + "neighbor_range": 1 + }, + "lattice": { + "type": "chain", + "length": null, + "boundary": "infinite", + "local_dim": 2, + "unit_cell": 2 + }, + "symmetry": { + "U1_Sz": true, + "parity": false, + "translation": true, + "target_sector": { + "total_Sz": 0 + } + }, + "ansatz": { + "family": "mps", + "algorithm": "vumps", + "variant": "two_site", + "initial_state": "neel", + "allow_bond_growth": true, + "target_precision": 1e-06 + } + }, + "capability": { + "capability_id": "tenpy.infinite_1d.vumps", + "maturity": "experimental", + "known_limitations": [ + "TeNPy UniformMPS is experimental.", + "Energy variance is backend-limited on this route.", + "The fixture does not count as a generated challenge problem." + ] + }, + "backend_binding": { + "schema_version": "tn-agent.backend-binding.v1", + "capability_id": "tenpy.infinite_1d.vumps", + "adapter_id": "tenpy.v1", + "backend_id": "tenpy", + "request_schema": "tn-agent.tenpy.infinite-xxz-vumps.v1", + "result_schema": "tn-agent.backend-result.v1" + }, + "numerics": { + "max_bond_dim": 32, + "cutoff": 1e-10, + "max_sweeps": 4, + "min_sweeps": 2, + "mixer": false, + "lanczos_maxiter": 8, + "checkpoint_every": 1, + "entropy_tolerance": 1e-05, + "finite_entanglement_fit": { + "enabled": true, + "min_chi": 16, + "max_chi": 32 + }, + "transfer_matrix": { + "compute": true, + "num_eigs": 4 + }, + "seed": 2468 + }, + "observables": [ + "energy", + "variance", + "entanglement_entropy", + "transfer_spectrum", + "central_charge_fit" + ], + "validators": [ + { + "id": "parse_consistency", + "policy": "required_pass", + "metric": null, + "operator": null, + "threshold": null + }, + { + "id": "convergence", + "policy": "required_pass", + "metric": "energy_drift", + "operator": "max", + "threshold": 0.0002 + }, + { + "id": "variance", + "policy": "backend_limited", + "metric": null, + "operator": null, + "threshold": null + }, + { + "id": "canonical_form", + "policy": "reported_only", + "metric": "canonical_residual", + "operator": null, + "threshold": null + }, + { + "id": "symmetry_check", + "policy": "reported_only", + "metric": "symmetry_residual", + "operator": null, + "threshold": null + }, + { + "id": "benchmark_compare", + "policy": "required_pass", + "metric": "benchmark_delta", + "operator": "max", + "threshold": 0.001 + }, + { + "id": "reproducibility", + "policy": "reported_only", + "metric": "reproduction_delta", + "operator": null, + "threshold": null + }, + { + "id": "artifact_completeness", + "policy": "required_pass", + "metric": "missing_artifacts", + "operator": "equals", + "threshold": 0 + } + ], + "acceptance": { + "mode": "all_required", + "required_validator_ids": [ + "parse_consistency", + "convergence", + "benchmark_compare", + "artifact_completeness" + ], + "allowed_backend_limited_ids": [ + "variance" + ], + "require_execution_success": true, + "reported_only_validator_ids": [ + "canonical_form", + "symmetry_check", + "reproducibility" + ] + }, + "provenance": { + "created_by": "WangTheoPhys", + "created_at": "2026-07-29T00:00:00Z", + "sources": [ + { + "kind": "repository_skill", + "uri": "skills/method-mps/SKILL.md", + "sha256": "sha256:297e1df7bb54e403edf86b4e5b683a6774d0db78673cd0fc76dc116332dbf088" + }, + { + "kind": "repository_skill", + "uri": "skills/using-tenpy/SKILL.md", + "sha256": "sha256:622b88619105cc0643bfde6d295377dd0c90ceee297119029362e47e80c14f5d" + } + ], + "generation_log_uri": "tracks/agent-kb/solutions/WangTheoPhys/README.md", + "human_gatekeeper_role": "Ratify the Hamiltonian, geometry, symmetry sector, target observable, and preregistered acceptance gate before execution." + }, + "reference": { + "observable": "energy", + "value": -0.4431471805599453, + "units": "J", + "normalization": "per-site", + "source": { + "kind": "registered_artifact", + "uri": "energy-reference.json", + "sha256": "sha256:6cacc4617490466b747b041270e6316ad90732c342441949bac118fc84652ef3" + } + } +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/gate.py b/tracks/agent-kb/solutions/WangTheoPhys/gate.py new file mode 100755 index 000000000..6428eeae4 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/gate.py @@ -0,0 +1,3359 @@ +#!/usr/bin/env python3 +"""Fail-closed, standard-library gate for the WangTheoPhys public capsule.""" + +from __future__ import annotations + +import argparse +import hashlib +import json +import math +import os +import re +import stat +import sys +from collections.abc import Mapping, Sequence +from datetime import datetime +from pathlib import Path, PurePosixPath +from typing import NoReturn + +EXPERIMENT_SCHEMA = "wangtheophys.tn-experiment.v1" +EVIDENCE_SCHEMA = "wangtheophys.tn-evidence.v1" +HEURISTIC_SCHEMA = "wangtheophys.tn-heuristic.v1" +BACKEND_RESULT_SCHEMA = "tn-agent.backend-result.v1" +MAX_DOCUMENT_BYTES = 4 * 1024 * 1024 +MAX_ARTIFACT_BYTES = 1024 * 1024 +MAX_ARTIFACT_TOTAL_BYTES = 8 * 1024 * 1024 +MAX_ARTIFACTS = 64 +MAX_LIBRARY_RECORDS = 10_000 +READ_CHUNK_BYTES = 64 * 1024 +TEAM_ROOT = Path(__file__).resolve().parent +REPOSITORY_ROOT = TEAM_ROOT.parents[3] +DIGEST_PATTERN = re.compile(r"^sha256:[0-9a-f]{64}$") +IDENTIFIER_PATTERN = re.compile(r"^[A-Za-z0-9][A-Za-z0-9_.:@-]{0,127}$") +TIMESTAMP_PATTERN = re.compile( + r"^[0-9]{4}-[0-9]{2}-[0-9]{2}T[0-9]{2}:[0-9]{2}:[0-9]{2}Z$" +) +SKILL_URI_PATTERN = re.compile(r"^skills/[a-z0-9][a-z0-9-]{0,63}/SKILL[.]md$") +AUDIT_URI_PATTERN = re.compile(r"^(?:docs|tests)/(?:[A-Za-z0-9._-]+/)*[A-Za-z0-9._-]+$") +GATE_REASON_CODES = frozenset( + { + "ACCEPTANCE_CONTRACT_INVALID", + "ARTIFACT_DIGEST_MISMATCH", + "ARTIFACT_IO_ERROR", + "ARTIFACT_LIMIT_EXCEEDED", + "ARTIFACT_TOO_LARGE", + "ARTIFACT_UNSAFE_PATH", + "BINDING_MISMATCH", + "CLI_USAGE_ERROR", + "DOCUMENT_DUPLICATE_KEY", + "DOCUMENT_INVALID_JSON", + "DOCUMENT_INVALID_UTF8", + "DOCUMENT_IO_ERROR", + "DOCUMENT_NONFINITE", + "DOCUMENT_NOT_CANONICAL", + "DOCUMENT_TOO_LARGE", + "DOCUMENT_UNSAFE_PATH", + "EVIDENCE_ARTIFACT_MISSING", + "EXECUTION_NOT_SUCCEEDED", + "EXPERIMENT_DIGEST_MISMATCH", + "INTERNAL_ERROR", + "LIBRARY_RECORD_INVALID", + "LIBRARY_RECORD_LIMIT", + "LIBRARY_SEQUENCE_INVALID", + "MISSING_FIELD", + "OBSERVABLE_SET_MISMATCH", + "OBSERVABLE_STATUS_INVALID", + "PROVENANCE_MISMATCH", + "RESULT_DIGEST_MISMATCH", + "SCIENTIFIC_EVIDENCE_UNATTESTED", + "SCHEMA_VERSION_UNSUPPORTED", + "SECURE_FILE_IO_UNAVAILABLE", + "TYPE_MISMATCH", + "UNKNOWN_FIELD", + "UNSUPPORTED_ROUTE", + "VALIDATOR_FAILED", + "VALIDATOR_POLICY_MISMATCH", + "VALIDATOR_SET_MISMATCH", + "VALIDATOR_STATUS_INVALID", + "VALIDATOR_THRESHOLD_FAILED", + "VALUE_INVALID", + } +) + +VALIDATOR_IDS = frozenset( + { + "parse_consistency", + "convergence", + "variance", + "canonical_form", + "symmetry_check", + "benchmark_compare", + "reproducibility", + "artifact_completeness", + } +) +REPORTED_ONLY_VALIDATOR_IDS = frozenset( + { + "variance", + "canonical_form", + "symmetry_check", + "reproducibility", + } +) + +VALIDATOR_RULES: dict[str, tuple[str | None, str | None, str]] = { + "parse_consistency": (None, None, "none"), + "convergence": ("energy_drift", "max", "nonnegative_number"), + "variance": ("variance", "max", "nonnegative_number"), + "canonical_form": ("canonical_residual", "max", "nonnegative_number"), + "symmetry_check": ("symmetry_residual", "max", "nonnegative_number"), + "benchmark_compare": ("benchmark_delta", "max", "nonnegative_number"), + "reproducibility": ("reproduction_delta", "max", "nonnegative_number"), + "artifact_completeness": ( + "missing_artifacts", + "equals", + "nonnegative_integer", + ), +} + +ROUTES: dict[str, dict[str, object]] = { + "tenpy.finite_1d.dmrg": { + "maturity": "stable", + "known_limitations": [ + "Only the finite spin-1/2 TFIM open-chain route is promoted.", + "The fixture does not count as a generated challenge problem.", + ], + "task_family": "ground_state_1d_finite", + "binding": { + "schema_version": "tn-agent.backend-binding.v1", + "capability_id": "tenpy.finite_1d.dmrg", + "adapter_id": "tenpy.v1", + "backend_id": "tenpy", + "request_schema": "tn-agent.tenpy.finite-tfim-dmrg.v1", + "result_schema": "tn-agent.backend-result.v1", + }, + "model_family": "tfim", + "operators": frozenset( + { + "sigma_x_i*sigma_x_i+1", + "sigma_z_i", + } + ), + "couplings": frozenset({"J", "g"}), + "boundary": "open", + "length": "positive", + "unit_cell": 1, + "symmetry": { + "U1_Sz": False, + "parity": False, + "translation": False, + "target_sector": None, + }, + "algorithm": "dmrg", + "variant": "two_site", + "initial_state": "all_z_plus", + "observables": frozenset( + { + "energy", + "variance", + "magnetization_z", + "entanglement_entropy", + "correlator_sigma_x", + "correlator_sz", + } + ), + "backend_limited_observables": frozenset(), + "observable_statuses": { + "energy": "measured", + "variance": "measured", + "magnetization_z": "measured", + "entanglement_entropy": "measured", + "correlator_sigma_x": "measured", + "correlator_sz": "measured", + }, + "backend_limited_validators": frozenset(), + }, + "tenpy.infinite_1d.vumps": { + "maturity": "experimental", + "known_limitations": [ + "TeNPy UniformMPS is experimental.", + "Energy variance is backend-limited on this route.", + "The fixture does not count as a generated challenge problem.", + ], + "task_family": "ground_state_1d_infinite", + "binding": { + "schema_version": "tn-agent.backend-binding.v1", + "capability_id": "tenpy.infinite_1d.vumps", + "adapter_id": "tenpy.v1", + "backend_id": "tenpy", + "request_schema": "tn-agent.tenpy.infinite-xxz-vumps.v1", + "result_schema": "tn-agent.backend-result.v1", + }, + "model_family": "xxz", + "operators": frozenset( + { + "Sx_i*Sx_i+1", + "Sy_i*Sy_i+1", + "Sz_i*Sz_i+1", + "Sz_i", + } + ), + "couplings": frozenset({"Jxy", "Delta", "h"}), + "boundary": "infinite", + "length": None, + "unit_cell": 2, + "symmetry": { + "U1_Sz": True, + "parity": False, + "translation": True, + "target_sector": {"total_Sz": 0}, + }, + "algorithm": "vumps", + "variant": "two_site", + "initial_state": "neel", + "observables": frozenset( + { + "energy", + "variance", + "magnetization_z", + "entanglement_entropy", + "transfer_spectrum", + "central_charge_fit", + } + ), + "backend_limited_observables": frozenset({"variance"}), + "observable_statuses": { + "energy": "measured", + "variance": "backend_limited", + "magnetization_z": "measured", + "entanglement_entropy": "measured", + "transfer_spectrum": "measured", + "central_charge_fit": "derived", + }, + "backend_limited_validators": frozenset({"variance"}), + }, +} + + +class GateError(Exception): + """A sanitized, stable rejection at a public contract boundary.""" + + def __init__( + self, + reason_code: str, + message: str, + *, + field: str | None = None, + exit_code: int = 2, + ) -> None: + super().__init__(message) + self.reason_code = reason_code + self.message = message + self.field = field + self.exit_code = exit_code + + def as_dict(self) -> dict[str, object]: + payload: dict[str, object] = { + "ok": False, + "reason_code": self.reason_code, + "message": self.message, + } + if self.exit_code == 3: + payload["accepted"] = False + if self.field is not None: + payload["field"] = self.field + return payload + + +def _fail( + reason_code: str, + message: str, + *, + field: str | None = None, + exit_code: int = 2, +) -> NoReturn: + if reason_code not in GATE_REASON_CODES: + raise RuntimeError("Unregistered public reason code") + raise GateError( + reason_code, + message, + field=field, + exit_code=exit_code, + ) + + +def canonical_digest(value: object) -> str: + """Return the stable SHA-256 of canonical UTF-8 JSON.""" + + try: + serialized = json.dumps( + value, + allow_nan=False, + ensure_ascii=False, + separators=(",", ":"), + sort_keys=True, + ).encode("utf-8") + except (TypeError, UnicodeError, ValueError): + _fail("DOCUMENT_NOT_CANONICAL", "Document is not canonical JSON") + return "sha256:" + hashlib.sha256(serialized).hexdigest() + + +def load_json_document(path: Path) -> object: + """Read one bounded regular file and parse strict JSON.""" + + raw = _read_regular_file(path, MAX_DOCUMENT_BYTES, artifact=False) + try: + text = raw.decode("utf-8", errors="strict") + except UnicodeDecodeError: + _fail("DOCUMENT_INVALID_UTF8", "Document is not valid UTF-8") + return _strict_json_loads(text) + + +def _strict_json_loads(text: str) -> object: + def reject_duplicates(pairs: list[tuple[str, object]]) -> dict[str, object]: + result: dict[str, object] = {} + for key, value in pairs: + if key in result: + _fail( + "DOCUMENT_DUPLICATE_KEY", + "Document contains a duplicate object key", + ) + result[key] = value + return result + + def reject_nonfinite(_value: str) -> NoReturn: + _fail("DOCUMENT_NONFINITE", "Document contains a non-finite number") + + def parse_float(value: str) -> float: + parsed = float(value) + if not math.isfinite(parsed): + _fail("DOCUMENT_NONFINITE", "Document contains a non-finite number") + return parsed + + try: + return json.loads( + text, + object_pairs_hook=reject_duplicates, + parse_constant=reject_nonfinite, + parse_float=parse_float, + ) + except GateError: + raise + except (json.JSONDecodeError, RecursionError, ValueError): + _fail("DOCUMENT_INVALID_JSON", "Document is not valid JSON") + + +def _read_regular_file(path: Path, limit: int, *, artifact: bool) -> bytes: + prefix = "ARTIFACT" if artifact else "DOCUMENT" + try: + before = path.lstat() + except (OSError, ValueError, UnicodeError): + _fail(f"{prefix}_IO_ERROR", f"{prefix.title()} cannot be read") + if stat.S_ISLNK(before.st_mode) or not stat.S_ISREG(before.st_mode): + _fail( + f"{prefix}_UNSAFE_PATH", + f"{prefix.title()} must be a regular non-symlink file", + ) + if before.st_size > limit: + _fail(f"{prefix}_TOO_LARGE", f"{prefix.title()} exceeds the byte limit") + if not hasattr(os, "O_NOFOLLOW"): + _fail( + "SECURE_FILE_IO_UNAVAILABLE", + "Secure no-follow reads are unavailable on this platform", + ) + flags = ( + os.O_RDONLY + | os.O_NOFOLLOW + | getattr(os, "O_CLOEXEC", 0) + | getattr(os, "O_NONBLOCK", 0) + ) + try: + descriptor = os.open(path, flags) + except (OSError, ValueError, UnicodeError): + _fail(f"{prefix}_IO_ERROR", f"{prefix.title()} cannot be opened safely") + try: + opened = os.fstat(descriptor) + if not stat.S_ISREG(opened.st_mode) or opened.st_size > limit: + _fail( + f"{prefix}_UNSAFE_PATH", + f"{prefix.title()} changed before it was opened", + ) + if not _same_file(before, opened): + _fail( + f"{prefix}_UNSAFE_PATH", + f"{prefix.title()} changed before it was opened", + ) + chunks: list[bytes] = [] + remaining = limit + 1 + while remaining: + chunk = os.read(descriptor, min(remaining, READ_CHUNK_BYTES)) + if not chunk: + break + chunks.append(chunk) + remaining -= len(chunk) + final = os.fstat(descriptor) + except OSError: + _fail(f"{prefix}_IO_ERROR", f"{prefix.title()} cannot be read safely") + finally: + os.close(descriptor) + raw = b"".join(chunks) + if len(raw) > limit: + _fail(f"{prefix}_TOO_LARGE", f"{prefix.title()} exceeds the byte limit") + if not _stable_file(opened, final): + _fail( + f"{prefix}_UNSAFE_PATH", + f"{prefix.title()} changed while it was read", + ) + return raw + + +def _same_file(left: os.stat_result, right: os.stat_result) -> bool: + return ( + stat.S_IFMT(left.st_mode) == stat.S_IFMT(right.st_mode) + and left.st_dev == right.st_dev + and left.st_ino == right.st_ino + ) + + +def _stable_file(left: os.stat_result, right: os.stat_result) -> bool: + return ( + _same_file(left, right) + and left.st_nlink == 1 + and right.st_nlink == 1 + and left.st_size == right.st_size + and getattr(left, "st_mtime_ns", left.st_mtime) + == getattr(right, "st_mtime_ns", right.st_mtime) + and getattr(left, "st_ctime_ns", left.st_ctime) + == getattr(right, "st_ctime_ns", right.st_ctime) + ) + + +def _object( + value: object, + expected_fields: frozenset[str], + field: str, +) -> dict[str, object]: + if type(value) is not dict: + _fail("TYPE_MISMATCH", "Expected a JSON object", field=field) + typed = value + unknown = sorted(set(typed) - expected_fields) + if unknown: + _fail( + "UNKNOWN_FIELD", + "Object contains an unknown field", + field=f"{field}.{unknown[0]}", + ) + missing = sorted(expected_fields - set(typed)) + if missing: + _fail( + "MISSING_FIELD", + "Object is missing a required field", + field=f"{field}.{missing[0]}", + ) + return typed + + +def _array(value: object, field: str, *, minimum: int = 0) -> list[object]: + if type(value) is not list: + _fail("TYPE_MISMATCH", "Expected a JSON array", field=field) + if len(value) < minimum: + _fail("VALUE_INVALID", "Array has too few items", field=field) + return value + + +def _string(value: object, field: str, *, identifier: bool = False) -> str: + if type(value) is not str: + _fail("TYPE_MISMATCH", "Expected a string", field=field) + if not value.strip(): + _fail("VALUE_INVALID", "String must not be empty", field=field) + if identifier and IDENTIFIER_PATTERN.fullmatch(value) is None: + _fail("VALUE_INVALID", "Identifier has an invalid shape", field=field) + return value + + +def _boolean(value: object, field: str) -> bool: + if type(value) is not bool: + _fail("TYPE_MISMATCH", "Expected a boolean", field=field) + return value + + +def _integer( + value: object, + field: str, + *, + minimum: int | None = None, +) -> int: + if type(value) is not int: + _fail("TYPE_MISMATCH", "Expected an integer", field=field) + if minimum is not None and value < minimum: + _fail("VALUE_INVALID", "Integer is below the allowed minimum", field=field) + return value + + +def _number( + value: object, + field: str, + *, + minimum: float | None = None, + maximum: float | None = None, +) -> float: + if type(value) not in {int, float}: + _fail("TYPE_MISMATCH", "Expected a finite number", field=field) + try: + result = float(value) + except (OverflowError, ValueError): + _fail( + "VALUE_INVALID", + "Number is outside the supported finite range", + field=field, + ) + if not math.isfinite(result): + _fail( + "DOCUMENT_NONFINITE", "Document contains a non-finite number", field=field + ) + if minimum is not None and result < minimum: + _fail("VALUE_INVALID", "Number is below the allowed minimum", field=field) + if maximum is not None and result > maximum: + _fail("VALUE_INVALID", "Number is above the allowed maximum", field=field) + return result + + +def _nullable_number(value: object, field: str) -> float | None: + return None if value is None else _number(value, field) + + +def _digest(value: object, field: str) -> str: + result = _string(value, field) + if DIGEST_PATTERN.fullmatch(result) is None: + _fail("VALUE_INVALID", "Expected a sha256: digest", field=field) + return result + + +def _timestamp(value: object, field: str) -> str: + result = _string(value, field) + if TIMESTAMP_PATTERN.fullmatch(result) is None: + _fail("VALUE_INVALID", "Expected a UTC RFC 3339 timestamp", field=field) + try: + datetime.fromisoformat(result[:-1] + "+00:00") + except ValueError: + _fail("VALUE_INVALID", "Expected a UTC RFC 3339 timestamp", field=field) + return result + + +def _unique_strings( + value: object, + field: str, + *, + minimum: int = 0, + identifiers: bool = False, +) -> list[str]: + items = _array(value, field, minimum=minimum) + result = [ + _string(item, f"{field}[{index}]", identifier=identifiers) + for index, item in enumerate(items) + ] + if len(result) != len(set(result)): + _fail("VALUE_INVALID", "Array values must be unique", field=field) + return result + + +def _expect(value: object, expected: object, field: str) -> None: + if value != expected: + _fail("UNSUPPORTED_ROUTE", "Value is outside the promoted route", field=field) + + +def _validate_binding(value: object, field: str) -> dict[str, object]: + binding = _object( + value, + frozenset( + { + "schema_version", + "capability_id", + "adapter_id", + "backend_id", + "request_schema", + "result_schema", + } + ), + field, + ) + _expect( + binding["schema_version"], + "tn-agent.backend-binding.v1", + f"{field}.schema_version", + ) + for name in ( + "capability_id", + "adapter_id", + "backend_id", + "request_schema", + "result_schema", + ): + _string(binding[name], f"{field}.{name}", identifier=True) + return binding + + +def validate_experiment(document: object) -> dict[str, object]: + """Validate one versioned experiment and its exact promoted route.""" + + root = _object( + document, + frozenset( + { + "schema_version", + "problem", + "physics", + "capability", + "backend_binding", + "numerics", + "observables", + "validators", + "acceptance", + "reference", + "provenance", + } + ), + "$", + ) + if root["schema_version"] != EXPERIMENT_SCHEMA: + _fail( + "SCHEMA_VERSION_UNSUPPORTED", + "Experiment schema version is unsupported", + field="$.schema_version", + ) + + problem = _object( + root["problem"], + frozenset( + { + "problem_id", + "title", + "research_question", + "novelty_claim", + "status", + "task_family", + } + ), + "$.problem", + ) + _string(problem["problem_id"], "$.problem.problem_id", identifier=True) + for name in ("title", "research_question", "novelty_claim"): + _string(problem[name], f"$.problem.{name}") + if problem["status"] not in {"candidate", "test_fixture"}: + _fail("VALUE_INVALID", "Problem status is invalid", field="$.problem.status") + task_family = _string( + problem["task_family"], "$.problem.task_family", identifier=True + ) + + physics = _object( + root["physics"], + frozenset({"model", "lattice", "symmetry", "ansatz"}), + "$.physics", + ) + model = _object( + physics["model"], + frozenset( + { + "representation", + "family", + "operators", + "couplings", + "onsite_terms", + "neighbor_range", + } + ), + "$.physics.model", + ) + _string(model["representation"], "$.physics.model.representation", identifier=True) + model_family = _string(model["family"], "$.physics.model.family", identifier=True) + operators = _unique_strings( + model["operators"], + "$.physics.model.operators", + minimum=1, + ) + couplings = _object( + model["couplings"], + frozenset(model["couplings"]) + if type(model["couplings"]) is dict + else frozenset(), + "$.physics.model.couplings", + ) + for name, value in couplings.items(): + _string(name, "$.physics.model.couplings.", identifier=True) + _number(value, f"$.physics.model.couplings.{name}") + onsite_terms = _array(model["onsite_terms"], "$.physics.model.onsite_terms") + _integer(model["neighbor_range"], "$.physics.model.neighbor_range", minimum=1) + + lattice = _object( + physics["lattice"], + frozenset({"type", "length", "boundary", "local_dim", "unit_cell"}), + "$.physics.lattice", + ) + _string(lattice["type"], "$.physics.lattice.type", identifier=True) + if lattice["length"] is not None: + _integer(lattice["length"], "$.physics.lattice.length", minimum=1) + _string(lattice["boundary"], "$.physics.lattice.boundary", identifier=True) + _integer(lattice["local_dim"], "$.physics.lattice.local_dim", minimum=1) + _integer(lattice["unit_cell"], "$.physics.lattice.unit_cell", minimum=1) + + symmetry = _object( + physics["symmetry"], + frozenset({"U1_Sz", "parity", "translation", "target_sector"}), + "$.physics.symmetry", + ) + for name in ("U1_Sz", "parity", "translation"): + _boolean(symmetry[name], f"$.physics.symmetry.{name}") + if symmetry["target_sector"] is not None: + target = _object( + symmetry["target_sector"], + frozenset({"total_Sz"}), + "$.physics.symmetry.target_sector", + ) + _integer(target["total_Sz"], "$.physics.symmetry.target_sector.total_Sz") + + ansatz = _object( + physics["ansatz"], + frozenset( + { + "family", + "algorithm", + "variant", + "initial_state", + "allow_bond_growth", + "target_precision", + } + ), + "$.physics.ansatz", + ) + for name in ("family", "algorithm", "variant", "initial_state"): + _string(ansatz[name], f"$.physics.ansatz.{name}", identifier=True) + _boolean(ansatz["allow_bond_growth"], "$.physics.ansatz.allow_bond_growth") + _number( + ansatz["target_precision"], + "$.physics.ansatz.target_precision", + minimum=0.0, + ) + if float(ansatz["target_precision"]) <= 0.0: + _fail( + "VALUE_INVALID", + "Target precision must be positive", + field="$.physics.ansatz.target_precision", + ) + + capability = _object( + root["capability"], + frozenset({"capability_id", "maturity", "known_limitations"}), + "$.capability", + ) + capability_id = _string( + capability["capability_id"], + "$.capability.capability_id", + identifier=True, + ) + route = ROUTES.get(capability_id) + if route is None: + _fail( + "UNSUPPORTED_ROUTE", + "Capability is not an executable promoted route", + field="$.capability.capability_id", + ) + _string(capability["maturity"], "$.capability.maturity", identifier=True) + _unique_strings( + capability["known_limitations"], + "$.capability.known_limitations", + minimum=1, + ) + binding = _validate_binding(root["backend_binding"], "$.backend_binding") + + numerics = _validate_numerics(root["numerics"], capability_id=capability_id) + observables = _unique_strings( + root["observables"], + "$.observables", + minimum=1, + identifiers=True, + ) + validators = _validate_validator_specs(root["validators"]) + acceptance = _validate_acceptance(root["acceptance"]) + reference = _validate_reference(root["reference"]) + _validate_provenance(root["provenance"]) + + _expect(task_family, route["task_family"], "$.problem.task_family") + _expect(capability["maturity"], route["maturity"], "$.capability.maturity") + _expect( + capability["known_limitations"], + route["known_limitations"], + "$.capability.known_limitations", + ) + _expect(binding, route["binding"], "$.backend_binding") + _expect(model["representation"], "operator_sum", "$.physics.model.representation") + _expect(model_family, route["model_family"], "$.physics.model.family") + if set(operators) != route["operators"] or len(operators) != len( + route["operators"] + ): + _fail( + "UNSUPPORTED_ROUTE", + "Operator set is outside the promoted route", + field="$.physics.model.operators", + ) + if set(couplings) != route["couplings"]: + _fail( + "UNSUPPORTED_ROUTE", + "Coupling set is outside the promoted route", + field="$.physics.model.couplings", + ) + if capability_id == "tenpy.infinite_1d.vumps" and couplings["Jxy"] != 1.0: + _fail( + "UNSUPPORTED_ROUTE", + "Standalone infinite XXZ is limited to Jxy=1", + field="$.physics.model.couplings.Jxy", + ) + _expect(onsite_terms, [], "$.physics.model.onsite_terms") + _expect(model["neighbor_range"], 1, "$.physics.model.neighbor_range") + _expect(lattice["type"], "chain", "$.physics.lattice.type") + _expect(lattice["boundary"], route["boundary"], "$.physics.lattice.boundary") + route_length = route["length"] + if route_length == "positive": + if type(lattice["length"]) is not int or lattice["length"] <= 0: + _fail( + "UNSUPPORTED_ROUTE", + "Finite route requires a positive chain length", + field="$.physics.lattice.length", + ) + else: + _expect(lattice["length"], route_length, "$.physics.lattice.length") + _expect(lattice["local_dim"], 2, "$.physics.lattice.local_dim") + _expect(lattice["unit_cell"], route["unit_cell"], "$.physics.lattice.unit_cell") + _expect(symmetry, route["symmetry"], "$.physics.symmetry") + _expect(ansatz["family"], "mps", "$.physics.ansatz.family") + _expect(ansatz["algorithm"], route["algorithm"], "$.physics.ansatz.algorithm") + _expect(ansatz["variant"], route["variant"], "$.physics.ansatz.variant") + _expect( + ansatz["initial_state"], + route["initial_state"], + "$.physics.ansatz.initial_state", + ) + _expect(ansatz["allow_bond_growth"], True, "$.physics.ansatz.allow_bond_growth") + + supported_observables = route["observables"] + if not set(observables).issubset(supported_observables): # type: ignore[arg-type] + _fail( + "UNSUPPORTED_ROUTE", + "Observable is outside the promoted route", + field="$.observables", + ) + if not {"energy", "variance"}.issubset(observables): + _fail( + "OBSERVABLE_SET_MISMATCH", + "Promoted acceptance requires energy and variance observables", + field="$.observables", + ) + backend_limited_validators = route["backend_limited_validators"] + expected_required = ( + VALIDATOR_IDS - REPORTED_ONLY_VALIDATOR_IDS - backend_limited_validators # type: ignore[operator] + ) + expected_limited = backend_limited_validators + expected_reported = REPORTED_ONLY_VALIDATOR_IDS - backend_limited_validators # type: ignore[operator] + actual_required = { + item["id"] for item in validators if item["policy"] == "required_pass" + } + actual_limited = { + item["id"] for item in validators if item["policy"] == "backend_limited" + } + actual_reported = { + item["id"] for item in validators if item["policy"] == "reported_only" + } + if ( + actual_required != expected_required + or actual_limited != expected_limited + or actual_reported != expected_reported + ): + _fail( + "VALIDATOR_POLICY_MISMATCH", + "Validator policy does not match the promoted route", + field="$.validators", + ) + if set(acceptance["required_validator_ids"]) != actual_required: + _fail( + "ACCEPTANCE_CONTRACT_INVALID", + "Acceptance required validators do not match validator policy", + field="$.acceptance.required_validator_ids", + ) + if set(acceptance["allowed_backend_limited_ids"]) != actual_limited: + _fail( + "ACCEPTANCE_CONTRACT_INVALID", + "Acceptance backend-limited validators do not match validator policy", + field="$.acceptance.allowed_backend_limited_ids", + ) + if set(acceptance["reported_only_validator_ids"]) != actual_reported: + _fail( + "ACCEPTANCE_CONTRACT_INVALID", + "Acceptance reported-only validators do not match validator policy", + field="$.acceptance.reported_only_validator_ids", + ) + expected_normalization = ( + "total" if capability_id == "tenpy.finite_1d.dmrg" else "per-site" + ) + if reference["normalization"] != expected_normalization: + _fail( + "VALUE_INVALID", + "Reference normalization does not match the promoted route", + field="$.reference.normalization", + ) + _validate_route_numerics(capability_id, numerics, observables) + return { + "ok": True, + "reason_code": "OK", + "schema_version": EXPERIMENT_SCHEMA, + "problem_id": problem["problem_id"], + "problem_status": problem["status"], + "capability_id": capability_id, + "experiment_digest": canonical_digest(root), + } + + +def _validate_numerics( + value: object, + *, + capability_id: str, +) -> dict[str, object]: + common_fields = frozenset( + { + "max_bond_dim", + "cutoff", + "max_sweeps", + "mixer", + "lanczos_maxiter", + "checkpoint_every", + "finite_entanglement_fit", + "transfer_matrix", + "seed", + } + ) + infinite_fields = frozenset({"min_sweeps", "entropy_tolerance"}) + numerics = _object( + value, + common_fields + | ( + infinite_fields + if capability_id == "tenpy.infinite_1d.vumps" + else frozenset() + ), + "$.numerics", + ) + for name in ( + "max_bond_dim", + "max_sweeps", + "lanczos_maxiter", + "checkpoint_every", + ): + _integer(numerics[name], f"$.numerics.{name}", minimum=1) + _integer(numerics["seed"], "$.numerics.seed", minimum=0) + _number(numerics["cutoff"], "$.numerics.cutoff", minimum=0.0) + _boolean(numerics["mixer"], "$.numerics.mixer") + if capability_id == "tenpy.infinite_1d.vumps": + _integer(numerics["min_sweeps"], "$.numerics.min_sweeps", minimum=1) + entropy = _number( + numerics["entropy_tolerance"], + "$.numerics.entropy_tolerance", + minimum=0.0, + ) + if entropy <= 0: + _fail( + "VALUE_INVALID", + "Entropy tolerance must be positive", + field="$.numerics.entropy_tolerance", + ) + fit = _object( + numerics["finite_entanglement_fit"], + frozenset({"enabled", "min_chi", "max_chi"}), + "$.numerics.finite_entanglement_fit", + ) + _boolean(fit["enabled"], "$.numerics.finite_entanglement_fit.enabled") + _integer(fit["min_chi"], "$.numerics.finite_entanglement_fit.min_chi", minimum=1) + _integer(fit["max_chi"], "$.numerics.finite_entanglement_fit.max_chi", minimum=1) + transfer = _object( + numerics["transfer_matrix"], + frozenset({"compute", "num_eigs"}), + "$.numerics.transfer_matrix", + ) + _boolean(transfer["compute"], "$.numerics.transfer_matrix.compute") + _integer(transfer["num_eigs"], "$.numerics.transfer_matrix.num_eigs", minimum=1) + if ( + capability_id == "tenpy.infinite_1d.vumps" + and numerics["min_sweeps"] > numerics["max_sweeps"] # type: ignore[operator] + ): + _fail( + "VALUE_INVALID", + "Minimum sweeps exceed maximum sweeps", + field="$.numerics.min_sweeps", + ) + if fit["min_chi"] > fit["max_chi"]: # type: ignore[operator] + _fail( + "VALUE_INVALID", + "Finite-entanglement chi range is reversed", + field="$.numerics.finite_entanglement_fit", + ) + return numerics + + +def _validate_route_numerics( + capability_id: str, + numerics: Mapping[str, object], + observables: Sequence[str], +) -> None: + fit = numerics["finite_entanglement_fit"] + transfer = numerics["transfer_matrix"] + assert isinstance(fit, dict) + assert isinstance(transfer, dict) + if capability_id == "tenpy.finite_1d.dmrg": + _expect(numerics["mixer"], True, "$.numerics.mixer") + _expect(fit["enabled"], False, "$.numerics.finite_entanglement_fit.enabled") + _expect(transfer["compute"], False, "$.numerics.transfer_matrix.compute") + else: + _expect(numerics["mixer"], False, "$.numerics.mixer") + if ( + type(numerics["entropy_tolerance"]) not in {int, float} + or float(numerics["entropy_tolerance"]) <= 0 + ): + _fail( + "UNSUPPORTED_ROUTE", + "Infinite route requires an explicit positive entropy tolerance", + field="$.numerics.entropy_tolerance", + ) + _expect(fit["enabled"], True, "$.numerics.finite_entanglement_fit.enabled") + _expect(transfer["compute"], True, "$.numerics.transfer_matrix.compute") + if fit["max_chi"] != numerics["max_bond_dim"]: + _fail( + "UNSUPPORTED_ROUTE", + "Fit maximum chi must equal the maximum bond dimension", + field="$.numerics.finite_entanglement_fit.max_chi", + ) + if "central_charge_fit" in observables and fit["min_chi"] == fit["max_chi"]: + _fail( + "UNSUPPORTED_ROUTE", + "Central-charge fitting requires at least two chi points", + field="$.numerics.finite_entanglement_fit", + ) + + +def _validate_validator_specs(value: object) -> list[dict[str, object]]: + items = _array(value, "$.validators", minimum=1) + validators: list[dict[str, object]] = [] + for index, raw in enumerate(items): + field = f"$.validators[{index}]" + item = _object( + raw, + frozenset({"id", "policy", "metric", "operator", "threshold"}), + field, + ) + validator_id = _string(item["id"], f"{field}.id", identifier=True) + if validator_id not in VALIDATOR_IDS: + _fail( + "UNSUPPORTED_ROUTE", + "Validator is outside the promoted route", + field=f"{field}.id", + ) + if item["policy"] not in { + "required_pass", + "reported_only", + "backend_limited", + }: + _fail( + "VALUE_INVALID", "Validator policy is invalid", field=f"{field}.policy" + ) + if item["metric"] is None: + if item["operator"] is not None or item["threshold"] is not None: + _fail( + "VALUE_INVALID", + "Metric-free validator must not define an operator or threshold", + field=field, + ) + else: + _string(item["metric"], f"{field}.metric", identifier=True) + if item["policy"] == "reported_only": + if item["operator"] is not None or item["threshold"] is not None: + _fail( + "VALIDATOR_POLICY_MISMATCH", + "Reported-only diagnostics cannot define acceptance thresholds", + field=field, + ) + elif item["operator"] not in {"max", "min", "equals"}: + _fail( + "VALUE_INVALID", + "Validator operator is invalid", + field=f"{field}.operator", + ) + else: + _number(item["threshold"], f"{field}.threshold") + if item["policy"] == "backend_limited" and item["metric"] is not None: + _fail( + "VALUE_INVALID", + "Backend-limited validator must not claim a metric threshold", + field=field, + ) + expected_metric, expected_operator, domain = VALIDATOR_RULES[validator_id] + if item["policy"] == "backend_limited": + expected_metric = None + expected_operator = None + elif item["policy"] == "reported_only": + expected_operator = None + if item["metric"] != expected_metric or item["operator"] != expected_operator: + _fail( + "VALIDATOR_POLICY_MISMATCH", + "Validator metric or operator does not match its public contract", + field=field, + ) + if item["policy"] == "required_pass" and expected_metric is not None: + if domain == "nonnegative_integer": + if type(item["threshold"]) is not int or item["threshold"] != 0: + _fail( + "VALIDATOR_POLICY_MISMATCH", + "Artifact completeness requires an exact integer zero threshold", + field=f"{field}.threshold", + ) + else: + _number(item["threshold"], f"{field}.threshold", minimum=0.0) + validators.append(item) + ids = [str(item["id"]) for item in validators] + if len(ids) != len(set(ids)): + _fail( + "VALUE_INVALID", + "Validator identifiers must be unique", + field="$.validators", + ) + if set(ids) != VALIDATOR_IDS: + _fail( + "VALIDATOR_SET_MISMATCH", + "Experiment must declare the complete public validator set", + field="$.validators", + ) + return validators + + +def _validate_acceptance(value: object) -> dict[str, object]: + acceptance = _object( + value, + frozenset( + { + "mode", + "required_validator_ids", + "allowed_backend_limited_ids", + "reported_only_validator_ids", + "require_execution_success", + } + ), + "$.acceptance", + ) + if acceptance["mode"] != "all_required": + _fail( + "ACCEPTANCE_CONTRACT_INVALID", + "Only all_required acceptance is supported", + field="$.acceptance.mode", + ) + acceptance["required_validator_ids"] = _unique_strings( + acceptance["required_validator_ids"], + "$.acceptance.required_validator_ids", + minimum=1, + identifiers=True, + ) + acceptance["allowed_backend_limited_ids"] = _unique_strings( + acceptance["allowed_backend_limited_ids"], + "$.acceptance.allowed_backend_limited_ids", + identifiers=True, + ) + acceptance["reported_only_validator_ids"] = _unique_strings( + acceptance["reported_only_validator_ids"], + "$.acceptance.reported_only_validator_ids", + identifiers=True, + ) + if set(acceptance["required_validator_ids"]) & set( + acceptance["allowed_backend_limited_ids"] + ): + _fail( + "ACCEPTANCE_CONTRACT_INVALID", + "Required and backend-limited validators overlap", + field="$.acceptance", + ) + policy_sets = ( + set(acceptance["required_validator_ids"]), + set(acceptance["allowed_backend_limited_ids"]), + set(acceptance["reported_only_validator_ids"]), + ) + if any( + policy_sets[index] & policy_sets[other] + for index in range(3) + for other in range(index + 1, 3) + ): + _fail( + "ACCEPTANCE_CONTRACT_INVALID", + "Required, reported-only, and backend-limited validator sets overlap", + field="$.acceptance", + ) + if ( + _boolean( + acceptance["require_execution_success"], + "$.acceptance.require_execution_success", + ) + is not True + ): + _fail( + "ACCEPTANCE_CONTRACT_INVALID", + "Successful execution must be required", + field="$.acceptance.require_execution_success", + ) + return acceptance + + +def _validate_reference(value: object) -> dict[str, object]: + reference = _object( + value, + frozenset({"observable", "value", "units", "normalization", "source"}), + "$.reference", + ) + _expect(reference["observable"], "energy", "$.reference.observable") + _number(reference["value"], "$.reference.value") + _expect(reference["units"], "J", "$.reference.units") + if reference["normalization"] not in {"total", "per-site"}: + _fail( + "VALUE_INVALID", + "Reference normalization is invalid", + field="$.reference.normalization", + ) + source = _object( + reference["source"], + frozenset({"kind", "uri", "sha256"}), + "$.reference.source", + ) + _expect(source["kind"], "registered_artifact", "$.reference.source.kind") + uri = _string(source["uri"], "$.reference.source.uri") + _validate_relative_path(uri, "$.reference.source.uri") + _digest(source["sha256"], "$.reference.source.sha256") + return reference + + +def _validate_provenance(value: object) -> None: + provenance = _object( + value, + frozenset( + { + "created_by", + "created_at", + "sources", + "generation_log_uri", + "human_gatekeeper_role", + } + ), + "$.provenance", + ) + _string(provenance["created_by"], "$.provenance.created_by") + _timestamp(provenance["created_at"], "$.provenance.created_at") + _string(provenance["generation_log_uri"], "$.provenance.generation_log_uri") + _string(provenance["human_gatekeeper_role"], "$.provenance.human_gatekeeper_role") + sources = _array(provenance["sources"], "$.provenance.sources", minimum=1) + identities: list[tuple[str, str, str]] = [] + for index, raw in enumerate(sources): + field = f"$.provenance.sources[{index}]" + source = _object(raw, frozenset({"kind", "uri", "sha256"}), field) + kind = _string(source["kind"], f"{field}.kind", identifier=True) + uri = _string(source["uri"], f"{field}.uri") + digest = _digest(source["sha256"], f"{field}.sha256") + identities.append((kind, uri, digest)) + if len(identities) != len(set(identities)): + _fail( + "VALUE_INVALID", + "Provenance sources must be unique", + field="$.provenance.sources", + ) + + +def _validate_evidence(document: object) -> dict[str, object]: + root = _object( + document, + frozenset( + { + "schema_version", + "experiment_digest", + "binding", + "execution", + "repeat_execution", + "artifacts", + "observables", + "validator_results", + "provenance", + "result_digest", + } + ), + "$", + ) + if root["schema_version"] != EVIDENCE_SCHEMA: + _fail( + "SCHEMA_VERSION_UNSUPPORTED", + "Evidence schema version is unsupported", + field="$.schema_version", + ) + _digest(root["experiment_digest"], "$.experiment_digest") + _validate_binding(root["binding"], "$.binding") + execution = _validate_execution(root["execution"], "$.execution") + repeat_execution = _validate_execution( + root["repeat_execution"], + "$.repeat_execution", + ) + if not isinstance(execution["execution_handle"], str) or not isinstance( + repeat_execution["execution_handle"], str + ): + _fail( + "PROVENANCE_MISMATCH", + "Successful primary and repeat evidence requires explicit handles", + field="$.execution", + ) + if execution["execution_handle"] == repeat_execution["execution_handle"]: + _fail( + "PROVENANCE_MISMATCH", + "Primary and repeat executions require distinct handles", + field="$.repeat_execution.execution_handle", + ) + + artifacts = _array(root["artifacts"], "$.artifacts", minimum=1) + if len(artifacts) > MAX_ARTIFACTS: + _fail("ARTIFACT_LIMIT_EXCEEDED", "Artifact count exceeds the limit") + artifact_paths: list[str] = [] + artifact_roles: list[str] = [] + expected_media_types = { + "backend_request": "application/json", + "backend_raw_result": "application/json", + "backend_result": "application/json", + "backend_repeat_raw_result": "application/json", + "backend_repeat_result": "application/json", + "energy_reference": "application/json", + "validator_evidence": "application/json", + "backend_stdout": "text/plain", + "backend_stderr": "text/plain", + "backend_repeat_stdout": "text/plain", + "backend_repeat_stderr": "text/plain", + } + for index, raw in enumerate(artifacts): + field = f"$.artifacts[{index}]" + item = _object( + raw, + frozenset( + { + "relative_path", + "digest", + "size_bytes", + "media_type", + "role", + } + ), + field, + ) + relative = _string(item["relative_path"], f"{field}.relative_path") + _validate_relative_path(relative, f"{field}.relative_path") + artifact_paths.append(relative) + _digest(item["digest"], f"{field}.digest") + _integer(item["size_bytes"], f"{field}.size_bytes", minimum=0) + media_type = _string(item["media_type"], f"{field}.media_type") + role = _string(item["role"], f"{field}.role", identifier=True) + if role not in expected_media_types: + _fail( + "VALUE_INVALID", + "Artifact role is not part of the public contract", + field=f"{field}.role", + ) + if media_type != expected_media_types[role]: + _fail( + "VALUE_INVALID", + "Artifact media type does not match its required role", + field=f"{field}.media_type", + ) + artifact_roles.append(role) + if len(artifact_paths) != len(set(artifact_paths)): + _fail("VALUE_INVALID", "Artifact paths must be unique", field="$.artifacts") + required_roles = set(expected_media_types) + if ( + len(artifact_roles) != len(required_roles) + or set(artifact_roles) != required_roles + ): + _fail( + "VALUE_INVALID", + "Evidence must register exactly one artifact for every required role", + field="$.artifacts", + ) + + observable_items = _array(root["observables"], "$.observables", minimum=1) + observable_names: list[str] = [] + for index, raw in enumerate(observable_items): + field = f"$.observables[{index}]" + item = _object( + raw, + frozenset({"name", "status", "evidence_digest"}), + field, + ) + observable_names.append(_string(item["name"], f"{field}.name", identifier=True)) + if item["status"] not in {"measured", "derived", "backend_limited"}: + _fail( + "VALUE_INVALID", "Observable status is invalid", field=f"{field}.status" + ) + _digest(item["evidence_digest"], f"{field}.evidence_digest") + if len(observable_names) != len(set(observable_names)): + _fail("VALUE_INVALID", "Observable names must be unique", field="$.observables") + + result_items = _array( + root["validator_results"], + "$.validator_results", + minimum=1, + ) + result_ids: list[str] = [] + for index, raw in enumerate(result_items): + field = f"$.validator_results[{index}]" + item = _object( + raw, + frozenset( + { + "id", + "status", + "reason_code", + "metric_value", + "evidence_digest", + } + ), + field, + ) + result_ids.append(_string(item["id"], f"{field}.id", identifier=True)) + if item["status"] not in { + "pass", + "fail", + "reported_only", + "backend_limited", + }: + _fail( + "VALUE_INVALID", "Validator status is invalid", field=f"{field}.status" + ) + _string(item["reason_code"], f"{field}.reason_code", identifier=True) + _nullable_number(item["metric_value"], f"{field}.metric_value") + _digest(item["evidence_digest"], f"{field}.evidence_digest") + if len(result_ids) != len(set(result_ids)): + _fail( + "VALUE_INVALID", + "Validator result identifiers must be unique", + field="$.validator_results", + ) + + provenance = _object( + root["provenance"], + frozenset( + { + "plan_id", + "request_digest", + "backend_result_digest", + "repeat_backend_result_digest", + "generated_by", + "generated_at", + } + ), + "$.provenance", + ) + for name in ( + "plan_id", + "request_digest", + "backend_result_digest", + "repeat_backend_result_digest", + ): + _digest(provenance[name], f"$.provenance.{name}") + _string(provenance["generated_by"], "$.provenance.generated_by") + _timestamp(provenance["generated_at"], "$.provenance.generated_at") + _digest(root["result_digest"], "$.result_digest") + return root + + +def _validate_execution(value: object, field: str) -> dict[str, object]: + execution = _object( + value, + frozenset( + { + "schema_version", + "status", + "return_code", + "execution_handle", + "retryable", + "stdout_digest", + "stderr_digest", + "stdout_truncated", + "stderr_truncated", + } + ), + field, + ) + if execution["schema_version"] != "tn-agent.execution-evidence.v1": + _fail( + "SCHEMA_VERSION_UNSUPPORTED", + "Execution evidence schema version is unsupported", + field=f"{field}.schema_version", + ) + if execution["status"] not in { + "succeeded", + "failed", + "timed_out", + "cancelled", + "not_executed", + }: + _fail("VALUE_INVALID", "Execution status is invalid", field=f"{field}.status") + if execution["return_code"] is not None: + _integer(execution["return_code"], f"{field}.return_code") + if execution["execution_handle"] is not None: + _string(execution["execution_handle"], f"{field}.execution_handle") + _boolean(execution["retryable"], f"{field}.retryable") + _digest(execution["stdout_digest"], f"{field}.stdout_digest") + _digest(execution["stderr_digest"], f"{field}.stderr_digest") + _boolean(execution["stdout_truncated"], f"{field}.stdout_truncated") + _boolean(execution["stderr_truncated"], f"{field}.stderr_truncated") + return execution + + +def _expected_plan_id( + experiment: Mapping[str, object], + experiment_digest: str, +) -> str: + return canonical_digest( + { + "schema_version": "tn-agent.execution-plan.v1", + "experiment_digest": experiment_digest, + "binding": experiment["backend_binding"], + "physics": experiment["physics"], + "numerics": experiment["numerics"], + "observables": experiment["observables"], + "validators": experiment["validators"], + } + ) + + +def _expected_request( + experiment: Mapping[str, object], + experiment_digest: str, + plan_id: str, +) -> dict[str, object]: + binding = experiment["backend_binding"] + physics = experiment["physics"] + numerics = experiment["numerics"] + validators = experiment["validators"] + assert isinstance(binding, dict) + assert isinstance(physics, dict) + assert isinstance(numerics, dict) + assert isinstance(validators, list) + model = physics["model"] + lattice = physics["lattice"] + symmetry = physics["symmetry"] + ansatz = physics["ansatz"] + fit = numerics["finite_entanglement_fit"] + transfer = numerics["transfer_matrix"] + assert isinstance(model, dict) + assert isinstance(lattice, dict) + assert isinstance(symmetry, dict) + assert isinstance(ansatz, dict) + assert isinstance(fit, dict) + assert isinstance(transfer, dict) + thresholds = { + str(item["id"]): item["threshold"] + for item in validators + if isinstance(item, dict) + } + acceptance = { + # The backend fields are execution/tuning inputs. Scientific acceptance + # is owned by this public gate, and raw-only diagnostics are never + # promoted merely because a worker reports a small value. + "require_all_validators": False, + "energy_drift_max": float(thresholds["convergence"]), + "variance_max": 0.0, + "canonical_residual_max": float(ansatz["target_precision"]), + "symmetry_residual_max": 0.0, + # The current worker schema requires this compatibility field, but the + # public gate deliberately gives repeat consistency no acceptance + # threshold until a trusted scheduler/registry receipt is available. + "reproducibility_max": 0.0, + } + max_bond_dim = int(numerics["max_bond_dim"]) + capability_id = str(binding["capability_id"]) + if capability_id == "tenpy.finite_1d.dmrg": + chi_schedule = [max_bond_dim] + effective_svd_min: float | None = float(numerics["cutoff"]) + engine_cutoff = float(numerics["cutoff"]) + minimum_sweeps = 0 + entropy_tolerance: float | None = None + lanczos_n_max = int(numerics["lanczos_maxiter"]) + else: + minimum_chi = int(fit["min_chi"]) + chi_schedule = [] + current = minimum_chi + while current < max_bond_dim: + chi_schedule.append(current) + current *= 2 + chi_schedule.append(max_bond_dim) + chi_schedule = sorted(set(chi_schedule)) + effective_svd_min = None + engine_cutoff = 0.0 + minimum_sweeps = int(numerics["min_sweeps"]) + entropy_tolerance = float(numerics["entropy_tolerance"]) + lanczos_n_max = int(numerics["lanczos_maxiter"]) * 4 + common: dict[str, object] = { + "schema_version": binding["request_schema"], + "capability_id": capability_id, + "plan_id": plan_id, + "model_family": model["family"], + "representation": model["representation"], + "operators": sorted(model["operators"]), # type: ignore[arg-type] + "boundary": lattice["boundary"], + "local_dimension": lattice["local_dim"], + "unit_cell": lattice["unit_cell"], + "algorithm": ansatz["algorithm"], + "variant": ansatz["variant"], + "active_sites": 2, + "initial_state": ansatz["initial_state"], + "allow_bond_growth": ansatz["allow_bond_growth"], + "target_precision": float(ansatz["target_precision"]), + "seed": numerics["seed"], + "numerics": { + "max_bond_dim": max_bond_dim, + "chi_schedule": chi_schedule, + "requested_cutoff": float(numerics["cutoff"]), + "effective_svd_min": effective_svd_min, + "engine_cutoff": engine_cutoff, + "max_sweeps": numerics["max_sweeps"], + "min_sweeps": minimum_sweeps, + "entropy_tolerance": entropy_tolerance, + "mixer": numerics["mixer"], + "lanczos_maxiter": numerics["lanczos_maxiter"], + "lanczos_n_max": lanczos_n_max, + "checkpoint_every": numerics["checkpoint_every"], + "diagonal_gauge_frequency": 0, + "check_overlap": False, + }, + "finite_entanglement_fit": { + **fit, + "entropy_statistic": "center", + }, + "transfer_matrix": transfer, + "acceptance": acceptance, + "requested_observables": sorted(experiment["observables"]), # type: ignore[arg-type] + } + couplings = model["couplings"] + assert isinstance(couplings, dict) + if capability_id == "tenpy.finite_1d.dmrg": + request = { + **common, + "length": lattice["length"], + "coupling_j": float(couplings["J"]), + "transverse_field_g": float(couplings["g"]), + "conserve": None, + "target_total_sz": None, + } + else: + target_sector = symmetry["target_sector"] + assert isinstance(target_sector, dict) + request = { + **common, + "coupling_jxy": float(couplings["Jxy"]), + "anisotropy_delta": float(couplings["Delta"]), + "field_h": float(couplings["h"]), + "conserve": "Sz", + "target_total_sz": target_sector["total_Sz"], + } + return request + + +def _strict_json_bytes(raw: bytes, label: str) -> object: + try: + return _strict_json_loads(raw.decode("utf-8", errors="strict")) + except UnicodeDecodeError: + _fail( + "DOCUMENT_INVALID_UTF8", + f"{label} is not valid UTF-8", + exit_code=3, + ) + + +def _validate_metric_domain( + validator_id: str, + value: object, + field: str, +) -> None: + _metric, _operator, domain = VALIDATOR_RULES[validator_id] + if domain == "none": + if value is not None: + _fail( + "VALIDATOR_STATUS_INVALID", + "Metric-free validator must have a null value", + field=field, + exit_code=3, + ) + elif domain == "nonnegative_integer": + if type(value) is not int or value < 0: + _fail( + "VALIDATOR_STATUS_INVALID", + "Artifact count must be an exact nonnegative integer", + field=field, + exit_code=3, + ) + else: + observed = _number(value, field) + if observed < 0.0: + _fail( + "VALIDATOR_STATUS_INVALID", + "Validator metric must be nonnegative", + field=field, + exit_code=3, + ) + + +def _validate_json_value(value: object, field: str) -> None: + if value is None or type(value) in {str, bool}: + return + if type(value) in {int, float}: + _number(value, field) + return + if type(value) is list: + for index, item in enumerate(value): + _validate_json_value(item, f"{field}[{index}]") + return + if type(value) is dict: + for key, item in value.items(): + _string(key, f"{field}.key") + _validate_json_value(item, f"{field}.{key}") + return + _fail("TYPE_MISMATCH", "Expected a JSON value", field=field) + + +def _validate_reconstructed_shape( + value: object, + expected: object, + field: str, +) -> None: + """Fail closed on shape and JSON domains before comparing canonical values.""" + + if type(expected) is dict: + expected_object = expected + actual = _object(value, frozenset(expected_object), field) + for name, expected_value in expected_object.items(): + _validate_reconstructed_shape( + actual[name], + expected_value, + f"{field}.{name}", + ) + return + if type(expected) is list: + actual = _array(value, field) + expected_list = expected + if len(actual) != len(expected_list): + _fail( + "VALUE_INVALID", + "Array length differs from the canonical reconstruction", + field=field, + exit_code=3, + ) + for index, (item, expected_item) in enumerate( + zip(actual, expected_list, strict=True) + ): + _validate_reconstructed_shape(item, expected_item, f"{field}[{index}]") + return + if expected is None: + if value is not None: + _fail( + "TYPE_MISMATCH", + "Expected null in the canonical reconstruction", + field=field, + exit_code=3, + ) + return + if type(expected) is str: + _string(value, field) + return + if type(expected) is bool: + _boolean(value, field) + return + if type(expected) is int: + _integer(value, field) + return + if type(expected) is float: + if type(value) is not float: + _fail( + "TYPE_MISMATCH", + "Expected an exact floating-point number", + field=field, + exit_code=3, + ) + _number(value, field) + return + _fail("TYPE_MISMATCH", "Unsupported reconstructed JSON value", field=field) + + +def _validate_raw_result( + raw: bytes, + *, + request: Mapping[str, object], + request_digest: str, + binding: Mapping[str, object], + route: Mapping[str, object], +) -> dict[str, object]: + document = _strict_json_bytes(raw, "Backend raw result") + result = _object( + document, + frozenset( + { + "schema_version", + "request_digest", + "plan_id", + "capability_id", + "backend_id", + "backend_version", + "adapter_id", + "adapter_version", + "status", + "environment", + "observables", + "convergence", + "warnings", + "known_limitations", + } + ), + "$raw", + ) + expected_identity = { + "schema_version": "tn-agent.tenpy.raw-result.v1", + "request_digest": request_digest, + "plan_id": request["plan_id"], + "capability_id": binding["capability_id"], + "backend_id": binding["backend_id"], + "backend_version": "1.1.0", + "adapter_id": binding["adapter_id"], + "adapter_version": "1.0.0", + "status": "succeeded", + } + for name, expected in expected_identity.items(): + if result[name] != expected: + _fail( + "PROVENANCE_MISMATCH", + "Raw result identity or execution status is inconsistent", + field=f"$raw.{name}", + exit_code=3, + ) + environment = _object( + result["environment"], + frozenset( + { + "schema_version", + "environment_digest", + "runtime_id", + "runtime_version", + "dependencies", + } + ), + "$raw.environment", + ) + if environment["schema_version"] != "tn-agent.tenpy.environment.v1": + _fail( + "SCHEMA_VERSION_UNSUPPORTED", + "Raw environment schema version is unsupported", + field="$raw.environment.schema_version", + exit_code=3, + ) + _digest(environment["environment_digest"], "$raw.environment.environment_digest") + _string(environment["runtime_id"], "$raw.environment.runtime_id", identifier=True) + _string(environment["runtime_version"], "$raw.environment.runtime_version") + dependencies = _object( + environment["dependencies"], + frozenset({"numpy", "tenpy"}), + "$raw.environment.dependencies", + ) + for name, version in dependencies.items(): + _string(version, f"$raw.environment.dependencies.{name}") + environment_content = { + key: value for key, value in environment.items() if key != "environment_digest" + } + if environment["environment_digest"] != canonical_digest(environment_content): + _fail( + "PROVENANCE_MISMATCH", + "Raw environment digest does not match canonical content", + field="$raw.environment.environment_digest", + exit_code=3, + ) + + observables = _object( + result["observables"], + frozenset(request["requested_observables"]), # type: ignore[arg-type] + "$raw.observables", + ) + statuses = route["observable_statuses"] + assert isinstance(statuses, dict) + for name, raw_observable in observables.items(): + field = f"$raw.observables.{name}" + observable = _object( + raw_observable, + frozenset({"status", "value", "units", "normalization", "reason"}), + field, + ) + expected_status = statuses[name] + if observable["status"] != expected_status: + _fail( + "OBSERVABLE_STATUS_INVALID", + "Raw observable status violates the exact route contract", + field=f"{field}.status", + exit_code=3, + ) + if observable["units"] is not None: + _string(observable["units"], f"{field}.units") + if observable["normalization"] is not None: + _string(observable["normalization"], f"{field}.normalization") + if expected_status == "backend_limited": + if observable["value"] is not None or not isinstance( + observable["reason"], str + ): + _fail( + "OBSERVABLE_STATUS_INVALID", + "Backend-limited raw evidence requires a reason and no value", + field=field, + exit_code=3, + ) + else: + if observable["value"] is None or observable["reason"] is not None: + _fail( + "OBSERVABLE_STATUS_INVALID", + "Measured or derived raw evidence requires a value and no reason", + field=field, + exit_code=3, + ) + _validate_json_value(observable["value"], f"{field}.value") + + convergence = _array(result["convergence"], "$raw.convergence", minimum=2) + prior_step: int | None = None + for index, raw_point in enumerate(convergence): + field = f"$raw.convergence[{index}]" + point = _object(raw_point, frozenset({"step", "metrics"}), field) + step = _integer(point["step"], f"{field}.step", minimum=0) + if prior_step is not None and step <= prior_step: + _fail( + "VALUE_INVALID", + "Raw convergence steps must be strictly increasing", + field=f"{field}.step", + exit_code=3, + ) + prior_step = step + metrics = point["metrics"] + if type(metrics) is not dict or not metrics: + _fail( + "TYPE_MISMATCH", + "Raw convergence metrics must be a nonempty object", + field=f"{field}.metrics", + exit_code=3, + ) + for name, value in metrics.items(): + _string(name, f"{field}.metrics.key", identifier=True) + if type(value) is not float: + _fail( + "TYPE_MISMATCH", + "Raw convergence metrics must be exact floats", + field=f"{field}.metrics.{name}", + exit_code=3, + ) + _number(value, f"{field}.metrics.{name}") + for name in ("warnings", "known_limitations"): + _unique_strings(result[name], f"$raw.{name}") + if result["known_limitations"] != route["known_limitations"]: + _fail( + "PROVENANCE_MISMATCH", + "Raw result limitations do not match the promoted route", + field="$raw.known_limitations", + exit_code=3, + ) + return result + + +def _reconstruct_backend_bundle( + *, + request_digest: str, + binding: Mapping[str, object], + raw: Mapping[str, object], + execution: Mapping[str, object], + raw_artifact: Mapping[str, object], +) -> dict[str, object]: + raw_observables = raw["observables"] + raw_convergence = raw["convergence"] + environment = raw["environment"] + assert isinstance(raw_observables, dict) + assert isinstance(raw_convergence, list) + assert isinstance(environment, dict) + observables = { + name: { + "schema_version": "tn-agent.observable-evidence.v1", + "name": name, + **value, + } + for name, value in sorted(raw_observables.items()) + if isinstance(value, dict) + } + convergence = [ + { + "schema_version": "tn-agent.convergence-point.v1", + "step": point["step"], + "metrics": point["metrics"], + } + for point in raw_convergence + if isinstance(point, dict) + ] + limited = [ + name + for name, value in sorted(raw_observables.items()) + if isinstance(value, dict) and value["status"] == "backend_limited" + ] + content: dict[str, object] = { + "schema_version": BACKEND_RESULT_SCHEMA, + "request_digest": request_digest, + "binding": binding, + "backend": { + "schema_version": "tn-agent.backend-identity.v1", + "backend_id": raw["backend_id"], + "backend_version": raw["backend_version"], + "adapter_id": raw["adapter_id"], + "adapter_version": raw["adapter_version"], + }, + "environment": { + "schema_version": "tn-agent.environment-identity.v1", + "environment_digest": environment["environment_digest"], + "runtime_id": environment["runtime_id"], + "runtime_version": environment["runtime_version"], + "dependencies": environment["dependencies"], + }, + "execution": execution, + "observables": observables, + "convergence": convergence, + "diagnostics": [], + "provenance": { + "schema_version": "tn-agent.provenance-evidence.v1", + "plan_id": raw["plan_id"], + "capability_id": binding["capability_id"], + "adapter_id": binding["adapter_id"], + "raw_result_relative": raw_artifact["relative_path"], + }, + "artifacts": [ + { + "schema_version": "tn-agent.result-artifact.v1", + "relative_path": raw_artifact["relative_path"], + "digest": raw_artifact["digest"], + "media_type": raw_artifact["media_type"], + "size_bytes": raw_artifact["size_bytes"], + } + ], + "warnings": raw["warnings"], + "known_limitations": raw["known_limitations"], + "backend_limited_fields": limited, + } + return {**content, "result_digest": canonical_digest(content)} + + +def _run_energy_metrics( + raw: Mapping[str, object], + *, + field: str, +) -> tuple[float, float]: + observables = raw["observables"] + convergence = raw["convergence"] + assert isinstance(observables, dict) + assert isinstance(convergence, list) + previous = convergence[-2] + latest = convergence[-1] + assert isinstance(previous, dict) + assert isinstance(latest, dict) + previous_metrics = previous["metrics"] + latest_metrics = latest["metrics"] + assert isinstance(previous_metrics, dict) + assert isinstance(latest_metrics, dict) + energy_observable = observables["energy"] + assert isinstance(energy_observable, dict) + energy = _number( + energy_observable["value"], + f"{field}.observables.energy.value", + ) + latest_energy = _number( + latest_metrics.get("energy"), + f"{field}.convergence[-1].metrics.energy", + ) + previous_energy = _number( + previous_metrics.get("energy"), + f"{field}.convergence[-2].metrics.energy", + ) + if energy != latest_energy: + _fail( + "VALIDATOR_STATUS_INVALID", + "Final convergence energy does not match the energy observable", + field=f"{field}.convergence[-1].metrics.energy", + exit_code=3, + ) + energy_drift = abs(latest_energy - previous_energy) + reported_energy_drift = latest_metrics.get("energy_drift") + if reported_energy_drift is not None: + reported = _number( + reported_energy_drift, + f"{field}.convergence[-1].metrics.energy_drift", + minimum=0.0, + ) + if reported != energy_drift: + _fail( + "VALIDATOR_STATUS_INVALID", + "Reported energy drift conflicts with the derived energy delta", + field=f"{field}.convergence[-1].metrics.energy_drift", + exit_code=3, + ) + return energy, energy_drift + + +def _validate_energy_reference( + raw: bytes, + *, + experiment: Mapping[str, object], +) -> dict[str, object]: + document = _strict_json_bytes(raw, "Energy reference") + reference = experiment["reference"] + binding = experiment["backend_binding"] + assert isinstance(reference, dict) + assert isinstance(binding, dict) + content = _object( + document, + frozenset( + { + "schema_version", + "reference_id", + "capability_id", + "physics_digest", + "observable", + "value", + "units", + "normalization", + "method", + "citation", + "result_digest", + } + ), + "$energy_reference", + ) + _expect( + content["schema_version"], + "wangtheophys.tn-energy-reference.v1", + "$energy_reference.schema_version", + ) + _string(content["reference_id"], "$energy_reference.reference_id", identifier=True) + _expect( + content["capability_id"], + binding["capability_id"], + "$energy_reference.capability_id", + ) + _expect( + content["physics_digest"], + canonical_digest(experiment["physics"]), + "$energy_reference.physics_digest", + ) + for name in ("observable", "value", "units", "normalization"): + _expect( + content[name], + reference[name], + f"$energy_reference.{name}", + ) + _number(content["value"], "$energy_reference.value") + _string(content["method"], "$energy_reference.method", identifier=True) + _string(content["citation"], "$energy_reference.citation") + semantic_content = { + key: value for key, value in content.items() if key != "result_digest" + } + if content["result_digest"] != canonical_digest(semantic_content): + _fail( + "RESULT_DIGEST_MISMATCH", + "Energy-reference semantic digest is inconsistent", + field="$energy_reference.result_digest", + exit_code=3, + ) + return content + + +def _derived_validator_results( + raw: Mapping[str, object], + *, + repeat_raw: Mapping[str, object], + reference: Mapping[str, object], +) -> list[dict[str, object]]: + observables = raw["observables"] + convergence = raw["convergence"] + assert isinstance(observables, dict) + assert isinstance(convergence, list) + latest = convergence[-1] + assert isinstance(latest, dict) + latest_metrics = latest["metrics"] + assert isinstance(latest_metrics, dict) + variance_observable = observables["variance"] + assert isinstance(variance_observable, dict) + energy, energy_drift = _run_energy_metrics(raw, field="$raw") + repeat_energy, _repeat_energy_drift = _run_energy_metrics( + repeat_raw, + field="$repeat_raw", + ) + + def residual(name: str) -> float: + value = _number( + latest_metrics.get(name), + f"$raw.convergence[-1].metrics.{name}", + minimum=0.0, + ) + return value + + reference_energy = _number( + reference["value"], + "$energy_reference.value", + ) + variance_value: object + variance_metric: object + if variance_observable["status"] == "backend_limited": + variance_metric = None + variance_value = None + else: + variance_metric = "variance" + variance_value = _number( + variance_observable["value"], + "$raw.observables.variance.value", + minimum=0.0, + ) + return [ + { + "id": "parse_consistency", + "metric": None, + "value": None, + "source": "gate.primary_and_repeat_bundle_reconstruction", + }, + { + "id": "convergence", + "metric": "energy_drift", + "value": energy_drift, + "source": "gate.primary_raw.convergence.energy_delta", + }, + { + "id": "variance", + "metric": variance_metric, + "value": variance_value, + "source": "reported.primary_raw.observables.variance", + }, + { + "id": "canonical_form", + "metric": "canonical_residual", + "value": residual("canonical_residual"), + "source": "reported.primary_raw.convergence.canonical_residual", + }, + { + "id": "symmetry_check", + "metric": "symmetry_residual", + "value": residual("symmetry_residual"), + "source": "reported.primary_raw.convergence.symmetry_residual", + }, + { + "id": "benchmark_compare", + "metric": "benchmark_delta", + "value": abs(energy - reference_energy), + "source": "gate.primary_energy_vs_preregistered_reference", + }, + { + "id": "reproducibility", + "metric": "reproduction_delta", + "value": abs(energy - repeat_energy), + "source": "reported.primary_energy_vs_repeat_raw", + }, + { + "id": "artifact_completeness", + "metric": "missing_artifacts", + "value": 0, + "source": "verified.required_artifact_roles", + }, + ] + + +def _validate_validator_evidence( + raw: bytes, + *, + request_digest: str, + backend_result_digest: str, + repeat_backend_result_digest: str, + reference_artifact_digest: str, + derived_results: list[dict[str, object]], +) -> dict[str, object]: + document = _strict_json_bytes(raw, "Validator evidence") + evidence = _object( + document, + frozenset( + { + "schema_version", + "request_digest", + "backend_result_digest", + "repeat_backend_result_digest", + "reference_artifact_digest", + "results", + "result_digest", + } + ), + "$validator_evidence", + ) + expected_content = { + "schema_version": "wangtheophys.tn-validator-evidence.v1", + "request_digest": request_digest, + "backend_result_digest": backend_result_digest, + "repeat_backend_result_digest": repeat_backend_result_digest, + "reference_artifact_digest": reference_artifact_digest, + "results": derived_results, + } + expected = { + **expected_content, + "result_digest": canonical_digest(expected_content), + } + if canonical_digest(evidence) != canonical_digest(expected): + _fail( + "VALIDATOR_STATUS_INVALID", + "Validator evidence does not match the gate-derived evaluation", + field="$validator_evidence", + exit_code=3, + ) + return expected + + +def _evaluate_artifact_chain( + *, + experiment: Mapping[str, object], + experiment_digest: str, + evidence: Mapping[str, object], + artifacts_by_role: Mapping[str, Mapping[str, object]], + verified_artifacts: Mapping[str, bytes], +) -> dict[str, object]: + binding = experiment["backend_binding"] + assert isinstance(binding, dict) + expected_plan_id = _expected_plan_id(experiment, experiment_digest) + expected_request = _expected_request( + experiment, + experiment_digest, + expected_plan_id, + ) + request_artifact = artifacts_by_role["backend_request"] + request_document = _strict_json_bytes( + verified_artifacts[str(request_artifact["digest"])], + "Backend request", + ) + if canonical_digest(request_document) != canonical_digest(expected_request): + _fail( + "PROVENANCE_MISMATCH", + "Backend request does not match the canonical promoted request", + field="$request", + exit_code=3, + ) + request_digest = canonical_digest(expected_request) + raw_artifact = artifacts_by_role["backend_raw_result"] + route = ROUTES[str(binding["capability_id"])] + raw_result = _validate_raw_result( + verified_artifacts[str(raw_artifact["digest"])], + request=expected_request, + request_digest=request_digest, + binding=binding, + route=route, + ) + execution = evidence["execution"] + repeat_execution = evidence["repeat_execution"] + assert isinstance(execution, dict) + assert isinstance(repeat_execution, dict) + stdout_artifact = artifacts_by_role["backend_stdout"] + stderr_artifact = artifacts_by_role["backend_stderr"] + if ( + execution["stdout_digest"] != stdout_artifact["digest"] + or execution["stderr_digest"] != stderr_artifact["digest"] + ): + _fail( + "PROVENANCE_MISMATCH", + "Execution stream identities do not match verified artifacts", + field="$.execution", + exit_code=3, + ) + repeat_stdout_artifact = artifacts_by_role["backend_repeat_stdout"] + repeat_stderr_artifact = artifacts_by_role["backend_repeat_stderr"] + if ( + repeat_execution["stdout_digest"] != repeat_stdout_artifact["digest"] + or repeat_execution["stderr_digest"] != repeat_stderr_artifact["digest"] + ): + _fail( + "PROVENANCE_MISMATCH", + "Repeat execution stream identities do not match verified artifacts", + field="$.repeat_execution", + exit_code=3, + ) + reconstructed = _reconstruct_backend_bundle( + request_digest=request_digest, + binding=binding, + raw=raw_result, + execution=execution, + raw_artifact=raw_artifact, + ) + backend_artifact = artifacts_by_role["backend_result"] + submitted = _strict_json_bytes( + verified_artifacts[str(backend_artifact["digest"])], + "Normalized backend result", + ) + _validate_reconstructed_shape(submitted, reconstructed, "$backend_result") + if canonical_digest(submitted) != canonical_digest(reconstructed): + _fail( + "RESULT_DIGEST_MISMATCH", + "Normalized backend result does not match gate reconstruction", + field="$backend_result", + exit_code=3, + ) + + repeat_raw_artifact = artifacts_by_role["backend_repeat_raw_result"] + if repeat_raw_artifact["digest"] == raw_artifact["digest"]: + _fail( + "PROVENANCE_MISMATCH", + "Repeat evidence cannot reuse the primary raw-result identity", + field="$.artifacts.backend_repeat_raw_result", + exit_code=3, + ) + repeat_raw_result = _validate_raw_result( + verified_artifacts[str(repeat_raw_artifact["digest"])], + request=expected_request, + request_digest=request_digest, + binding=binding, + route=route, + ) + repeat_reconstructed = _reconstruct_backend_bundle( + request_digest=request_digest, + binding=binding, + raw=repeat_raw_result, + execution=repeat_execution, + raw_artifact=repeat_raw_artifact, + ) + repeat_backend_artifact = artifacts_by_role["backend_repeat_result"] + repeat_submitted = _strict_json_bytes( + verified_artifacts[str(repeat_backend_artifact["digest"])], + "Repeat normalized backend result", + ) + _validate_reconstructed_shape( + repeat_submitted, + repeat_reconstructed, + "$repeat_backend_result", + ) + if canonical_digest(repeat_submitted) != canonical_digest(repeat_reconstructed): + _fail( + "RESULT_DIGEST_MISMATCH", + "Repeat normalized result does not match gate reconstruction", + field="$repeat_backend_result", + exit_code=3, + ) + + reference_artifact = artifacts_by_role["energy_reference"] + experiment_reference = experiment["reference"] + assert isinstance(experiment_reference, dict) + reference_source = experiment_reference["source"] + assert isinstance(reference_source, dict) + if ( + reference_source["uri"] != reference_artifact["relative_path"] + or reference_source["sha256"] != reference_artifact["digest"] + ): + _fail( + "PROVENANCE_MISMATCH", + "Energy-reference artifact does not match preregistered source identity", + field="$.reference.source", + exit_code=3, + ) + energy_reference = _validate_energy_reference( + verified_artifacts[str(reference_artifact["digest"])], + experiment=experiment, + ) + + derived_results = _derived_validator_results( + raw_result, + repeat_raw=repeat_raw_result, + reference=energy_reference, + ) + validator_artifact = artifacts_by_role["validator_evidence"] + validator_evidence = _validate_validator_evidence( + verified_artifacts[str(validator_artifact["digest"])], + request_digest=request_digest, + backend_result_digest=str(reconstructed["result_digest"]), + repeat_backend_result_digest=str(repeat_reconstructed["result_digest"]), + reference_artifact_digest=str(reference_artifact["digest"]), + derived_results=derived_results, + ) + provenance = evidence["provenance"] + assert isinstance(provenance, dict) + if ( + provenance["plan_id"] != expected_plan_id + or provenance["request_digest"] != request_digest + or provenance["backend_result_digest"] != backend_artifact["digest"] + or provenance["repeat_backend_result_digest"] + != repeat_backend_artifact["digest"] + ): + _fail( + "PROVENANCE_MISMATCH", + "Evidence provenance does not match the reconstructed artifact chains", + field="$.provenance", + exit_code=3, + ) + return { + "observables": reconstructed["observables"], + "validator_metrics": {str(item["id"]): item for item in derived_results}, + "backend_artifact_digest": backend_artifact["digest"], + "validator_artifact_digest": validator_artifact["digest"], + "backend_result_digest": reconstructed["result_digest"], + "validator_result_digest": validator_evidence["result_digest"], + } + + +def evaluate( + experiment: object, + evidence: object, + *, + artifact_root: Path, +) -> dict[str, object]: + """Evaluate normalized evidence against the preregistered experiment.""" + + experiment_summary = validate_experiment(experiment) + if experiment_summary["problem_status"] == "candidate": + _fail( + "SCIENTIFIC_EVIDENCE_UNATTESTED", + "Candidate evidence lacks a trusted execution or state certificate", + field="$.problem.status", + exit_code=3, + ) + experiment_root = experiment + assert isinstance(experiment_root, dict) + evidence_root = _validate_evidence(evidence) + if evidence_root["experiment_digest"] != experiment_summary["experiment_digest"]: + _fail( + "EXPERIMENT_DIGEST_MISMATCH", + "Evidence does not bind the canonical experiment", + field="$.experiment_digest", + exit_code=3, + ) + if evidence_root["binding"] != experiment_root["backend_binding"]: + _fail( + "BINDING_MISMATCH", + "Evidence binding does not match the experiment", + field="$.binding", + exit_code=3, + ) + digest_content = { + key: value for key, value in evidence_root.items() if key != "result_digest" + } + if evidence_root["result_digest"] != canonical_digest(digest_content): + _fail( + "RESULT_DIGEST_MISMATCH", + "Evidence result digest does not match canonical content", + field="$.result_digest", + exit_code=3, + ) + + acceptance = experiment_root["acceptance"] + execution = evidence_root["execution"] + repeat_execution = evidence_root["repeat_execution"] + assert isinstance(acceptance, dict) + assert isinstance(execution, dict) + assert isinstance(repeat_execution, dict) + if acceptance["require_execution_success"]: + for label, execution_record in ( + ("Primary", execution), + ("Repeat", repeat_execution), + ): + if not ( + execution_record["status"] == "succeeded" + and execution_record["return_code"] == 0 + and execution_record["retryable"] is False + ): + _fail( + "EXECUTION_NOT_SUCCEEDED", + f"{label} execution evidence is not a successful terminal result", + exit_code=3, + ) + + verified_artifacts = _verify_artifacts( + evidence_root["artifacts"], + artifact_root, + ) + artifacts = evidence_root["artifacts"] + assert isinstance(artifacts, list) + artifacts_by_role = { + str(item["role"]): item for item in artifacts if isinstance(item, dict) + } + evaluated_chain = _evaluate_artifact_chain( + experiment=experiment_root, + experiment_digest=str(experiment_summary["experiment_digest"]), + evidence=evidence_root, + artifacts_by_role=artifacts_by_role, + verified_artifacts=verified_artifacts, + ) + normalized_observables = evaluated_chain["observables"] + normalized_metrics = evaluated_chain["validator_metrics"] + backend_artifact_digest = evaluated_chain["backend_artifact_digest"] + validator_artifact_digest = evaluated_chain["validator_artifact_digest"] + assert isinstance(normalized_observables, dict) + assert isinstance(normalized_metrics, dict) + assert isinstance(backend_artifact_digest, str) + assert isinstance(validator_artifact_digest, str) + + expected_observables = experiment_root["observables"] + observable_results = evidence_root["observables"] + assert isinstance(expected_observables, list) + assert isinstance(observable_results, list) + actual_observables = { + str(item["name"]): item for item in observable_results if isinstance(item, dict) + } + if set(actual_observables) != set(expected_observables): + _fail( + "OBSERVABLE_SET_MISMATCH", + "Evidence observable set does not match the experiment", + exit_code=3, + ) + capability = experiment_root["capability"] + assert isinstance(capability, dict) + route = ROUTES[str(capability["capability_id"])] + observable_statuses = route["observable_statuses"] + assert isinstance(observable_statuses, dict) + for name, item in actual_observables.items(): + expected_status = observable_statuses[name] + normalized_item = normalized_observables[name] + assert isinstance(normalized_item, dict) + if ( + item["status"] != expected_status + or item["status"] != normalized_item["status"] + ): + _fail( + "OBSERVABLE_STATUS_INVALID", + "Observable status violates the route contract", + field=f"$.observables.{name}", + exit_code=3, + ) + if item["evidence_digest"] != backend_artifact_digest: + _fail( + "EVIDENCE_ARTIFACT_MISSING", + "Observable evidence must identify the normalized backend result", + field=f"$.observables.{name}.evidence_digest", + exit_code=3, + ) + + validator_specs = experiment_root["validators"] + validator_results = evidence_root["validator_results"] + assert isinstance(validator_specs, list) + assert isinstance(validator_results, list) + spec_by_id = { + str(item["id"]): item for item in validator_specs if isinstance(item, dict) + } + result_by_id = { + str(item["id"]): item for item in validator_results if isinstance(item, dict) + } + if set(result_by_id) != set(spec_by_id): + _fail( + "VALIDATOR_SET_MISMATCH", + "Validator result set does not match the experiment", + exit_code=3, + ) + for validator_id, spec in spec_by_id.items(): + result = result_by_id[validator_id] + normalized_metric = normalized_metrics[validator_id] + assert isinstance(normalized_metric, dict) + if result["evidence_digest"] != validator_artifact_digest: + _fail( + "EVIDENCE_ARTIFACT_MISSING", + "Validator result must identify the separate validator evidence", + field=f"$.validator_results.{validator_id}.evidence_digest", + exit_code=3, + ) + if spec["policy"] == "backend_limited": + if not ( + result["status"] == "backend_limited" + and result["reason_code"] == "BACKEND_LIMITED" + and result["metric_value"] is None + ): + _fail( + "VALIDATOR_STATUS_INVALID", + "Backend-limited validator evidence is inconsistent", + field=f"$.validator_results.{validator_id}", + exit_code=3, + ) + continue + if spec["policy"] == "reported_only": + _validate_metric_domain( + validator_id, + result["metric_value"], + f"$.validator_results.{validator_id}.metric_value", + ) + if ( + type(result["metric_value"]) is not type(normalized_metric["value"]) + or result["metric_value"] != normalized_metric["value"] + or result["status"] != "reported_only" + or result["reason_code"] != "REPORTED_ONLY" + ): + _fail( + "VALIDATOR_STATUS_INVALID", + "Reported-only diagnostic does not match the bound raw report", + field=f"$.validator_results.{validator_id}", + exit_code=3, + ) + continue + _validate_metric_domain( + validator_id, + result["metric_value"], + f"$.validator_results.{validator_id}.metric_value", + ) + if ( + type(result["metric_value"]) is not type(normalized_metric["value"]) + or result["metric_value"] != normalized_metric["value"] + ): + _fail( + "VALIDATOR_STATUS_INVALID", + "Validator metric does not match the gate-derived evaluation", + field=f"$.validator_results.{validator_id}.metric_value", + exit_code=3, + ) + if not ( + result["status"] == "pass" and result["reason_code"] == "VALIDATOR_PASS" + ): + _fail( + "VALIDATOR_FAILED", + "A required validator did not pass", + field=f"$.validator_results.{validator_id}", + exit_code=3, + ) + _evaluate_threshold(validator_id, spec, result) + + return { + "ok": True, + "accepted": True, + "reason_code": "ACCEPTANCE_PASSED", + "problem_id": experiment_summary["problem_id"], + "problem_status": experiment_summary["problem_status"], + "capability_id": experiment_summary["capability_id"], + "experiment_digest": experiment_summary["experiment_digest"], + "result_digest": evidence_root["result_digest"], + "verified_artifacts": len(artifacts), + } + + +def _evaluate_threshold( + validator_id: str, + spec: Mapping[str, object], + result: Mapping[str, object], +) -> None: + metric = spec["metric"] + value = result["metric_value"] + if metric is None: + if value is not None: + _fail( + "VALIDATOR_STATUS_INVALID", + "Metric-free validator supplied a metric value", + field=f"$.validator_results.{validator_id}.metric_value", + exit_code=3, + ) + return + if type(value) not in {int, float}: + _fail( + "VALIDATOR_STATUS_INVALID", + "Validator metric value is missing", + field=f"$.validator_results.{validator_id}.metric_value", + exit_code=3, + ) + _validate_metric_domain( + validator_id, + value, + f"$.validator_results.{validator_id}.metric_value", + ) + threshold = float(spec["threshold"]) # type: ignore[arg-type] + observed = float(value) + operator = spec["operator"] + passed = ( + (operator == "max" and observed <= threshold) + or (operator == "min" and observed >= threshold) + or (operator == "equals" and observed == threshold) + ) + if not passed: + _fail( + "VALIDATOR_THRESHOLD_FAILED", + "Validator metric violates its preregistered threshold", + field=f"$.validator_results.{validator_id}.metric_value", + exit_code=3, + ) + + +def _verify_artifacts(value: object, artifact_root: Path) -> dict[str, bytes]: + artifacts = _array(value, "$.artifacts", minimum=1) + root_descriptor = _open_artifact_root(artifact_root) + total = 0 + artifacts_by_digest: dict[str, bytes] = {} + try: + for item in artifacts: + assert isinstance(item, dict) + relative = str(item["relative_path"]) + _validate_relative_path(relative, "$.artifacts.relative_path") + declared_size = int(item["size_bytes"]) + if declared_size > MAX_ARTIFACT_BYTES: + _fail( + "ARTIFACT_LIMIT_EXCEEDED", + "Declared artifact size exceeds the limit", + exit_code=3, + ) + raw = _read_artifact_file( + root_descriptor, + relative, + MAX_ARTIFACT_BYTES, + ) + total += len(raw) + if total > MAX_ARTIFACT_TOTAL_BYTES: + _fail( + "ARTIFACT_LIMIT_EXCEEDED", + "Artifact aggregate size exceeds the limit", + exit_code=3, + ) + observed = "sha256:" + hashlib.sha256(raw).hexdigest() + if len(raw) != declared_size or observed != item["digest"]: + _fail( + "ARTIFACT_DIGEST_MISMATCH", + "Artifact bytes do not match the declared identity", + exit_code=3, + ) + artifacts_by_digest[observed] = raw + finally: + os.close(root_descriptor) + return artifacts_by_digest + + +def _validate_relative_path(value: str, field: str) -> None: + path = PurePosixPath(value) + if ( + not value + or value == "." + or not path.parts + or path.is_absolute() + or value != path.as_posix() + or "\\" in value + or value.startswith("/") + or any( + ord(character) < 0x20 + or ord(character) == 0x7F + or 0xD800 <= ord(character) <= 0xDFFF + for character in value + ) + or any(part in {"", ".", ".."} for part in path.parts) + or (len(value) >= 2 and value[0].isalpha() and value[1] == ":") + ): + _fail( + "ARTIFACT_UNSAFE_PATH", + "Artifact path must be normalized and relative", + field=field, + ) + + +def _open_artifact_root(root: Path) -> int: + if ( + not hasattr(os, "O_NOFOLLOW") + or not hasattr(os, "O_DIRECTORY") + or os.open not in os.supports_dir_fd + ): + _fail( + "SECURE_FILE_IO_UNAVAILABLE", + "Secure component-wise artifact reads are unavailable", + exit_code=3, + ) + try: + before = root.lstat() + except (OSError, ValueError, UnicodeError): + _fail("ARTIFACT_IO_ERROR", "Artifact root cannot be inspected", exit_code=3) + if stat.S_ISLNK(before.st_mode) or not stat.S_ISDIR(before.st_mode): + _fail( + "ARTIFACT_UNSAFE_PATH", + "Artifact root must be a non-symlink directory", + exit_code=3, + ) + flags = os.O_RDONLY | os.O_NOFOLLOW | os.O_DIRECTORY | getattr(os, "O_CLOEXEC", 0) + descriptor: int | None = None + try: + descriptor = os.open(root, flags) + opened = os.fstat(descriptor) + except (OSError, ValueError, UnicodeError): + if descriptor is not None: + os.close(descriptor) + _fail("ARTIFACT_IO_ERROR", "Artifact root cannot be opened safely", exit_code=3) + assert descriptor is not None + if not stat.S_ISDIR(opened.st_mode) or not _same_file(before, opened): + os.close(descriptor) + _fail( + "ARTIFACT_UNSAFE_PATH", + "Artifact root changed before it was opened", + exit_code=3, + ) + return descriptor + + +def _read_artifact_file(root: int | Path, relative: str, limit: int) -> bytes: + """Open every artifact path component relative to directory descriptors.""" + + _validate_relative_path(relative, "$.artifacts.relative_path") + if ( + not hasattr(os, "O_NOFOLLOW") + or not hasattr(os, "O_DIRECTORY") + or os.open not in os.supports_dir_fd + ): + _fail( + "SECURE_FILE_IO_UNAVAILABLE", + "Secure component-wise artifact reads are unavailable", + exit_code=3, + ) + directory_descriptors: list[int] = [] + file_descriptor: int | None = None + owned_root: int | None = None + root_descriptor = root + if isinstance(root, Path): + owned_root = _open_artifact_root(root) + root_descriptor = owned_root + flags = ( + os.O_RDONLY + | os.O_NOFOLLOW + | getattr(os, "O_CLOEXEC", 0) + | getattr(os, "O_NONBLOCK", 0) + ) + try: + current = os.dup(root_descriptor) + directory_descriptors.append(current) + components = PurePosixPath(relative).parts + for component in components[:-1]: + current = os.open( + component, + flags | os.O_DIRECTORY, + dir_fd=current, + ) + directory_descriptors.append(current) + file_descriptor = os.open( + components[-1], + flags, + dir_fd=current, + ) + opened = os.fstat(file_descriptor) + if not stat.S_ISREG(opened.st_mode) or opened.st_nlink != 1: + _fail( + "ARTIFACT_UNSAFE_PATH", + "Artifact must be a regular single-link non-symlink file", + exit_code=3, + ) + if opened.st_size > limit: + _fail( + "ARTIFACT_TOO_LARGE", + "Artifact exceeds the byte limit", + exit_code=3, + ) + chunks: list[bytes] = [] + remaining = limit + 1 + while remaining: + chunk = os.read(file_descriptor, min(remaining, READ_CHUNK_BYTES)) + if not chunk: + break + chunks.append(chunk) + remaining -= len(chunk) + final = os.fstat(file_descriptor) + except GateError: + raise + except (OSError, ValueError, UnicodeError): + _fail("ARTIFACT_IO_ERROR", "Artifact cannot be opened safely", exit_code=3) + finally: + if file_descriptor is not None: + os.close(file_descriptor) + for descriptor in reversed(directory_descriptors): + os.close(descriptor) + if owned_root is not None: + os.close(owned_root) + raw = b"".join(chunks) + if len(raw) > limit: + _fail( + "ARTIFACT_TOO_LARGE", + "Artifact exceeds the byte limit", + exit_code=3, + ) + if not _stable_file(opened, final): + _fail( + "ARTIFACT_UNSAFE_PATH", + "Artifact changed while it was read", + exit_code=3, + ) + return raw + + +def validate_library(path: Path) -> dict[str, object]: + """Validate an append-only JSONL heuristics library.""" + + raw = _read_regular_file(path, MAX_DOCUMENT_BYTES, artifact=False) + try: + text = raw.decode("utf-8", errors="strict") + except UnicodeDecodeError: + _fail("DOCUMENT_INVALID_UTF8", "Library is not valid UTF-8") + lines = text.splitlines() + if not lines or any(not line.strip() for line in lines): + _fail( + "LIBRARY_RECORD_INVALID", + "Library must contain one non-empty JSON object per line", + ) + if len(lines) > MAX_LIBRARY_RECORDS: + _fail("LIBRARY_RECORD_LIMIT", "Library record count exceeds the limit") + records: list[dict[str, object]] = [] + for index, line in enumerate(lines): + value = _strict_json_loads(line) + try: + record = _validate_heuristic_record(value, index) + except GateError as error: + if error.reason_code.startswith("DOCUMENT_"): + raise + _fail( + "LIBRARY_RECORD_INVALID", + "Library contains an invalid heuristic record", + field=f"$line[{index + 1}]", + ) + records.append(record) + _validate_library_sequence(records) + return { + "ok": True, + "reason_code": "OK", + "schema_version": HEURISTIC_SCHEMA, + "records": len(records), + "heuristics": len({str(record["heuristic_id"]) for record in records}), + "library_digest": canonical_digest(records), + } + + +def _validate_heuristic_record(value: object, index: int) -> dict[str, object]: + field = f"$line[{index + 1}]" + record = _object( + value, + frozenset( + { + "schema_version", + "record_id", + "heuristic_id", + "revision", + "recorded_at", + "applies_to", + "claim", + "action", + "source", + "evidence", + "confidence", + "contradicts", + "supersedes", + "claim_status", + } + ), + field, + ) + if record["schema_version"] != HEURISTIC_SCHEMA: + _fail( + "SCHEMA_VERSION_UNSUPPORTED", + "Heuristic schema version is unsupported", + field=f"{field}.schema_version", + ) + record_id = _string(record["record_id"], f"{field}.record_id", identifier=True) + heuristic_id = _string( + record["heuristic_id"], + f"{field}.heuristic_id", + identifier=True, + ) + revision = _integer(record["revision"], f"{field}.revision", minimum=1) + if record_id != f"{heuristic_id}@{revision}": + _fail( + "VALUE_INVALID", + "Record identity must be heuristic_id@revision", + field=f"{field}.record_id", + ) + _timestamp(record["recorded_at"], f"{field}.recorded_at") + applies_to = _unique_strings( + record["applies_to"], + f"{field}.applies_to", + minimum=1, + identifiers=True, + ) + if any(route != "*" and route not in ROUTES for route in applies_to): + _fail( + "VALUE_INVALID", + "Heuristic names an unknown route", + field=f"{field}.applies_to", + ) + _string(record["claim"], f"{field}.claim") + _string(record["action"], f"{field}.action") + source = _object( + record["source"], + frozenset({"kind", "uri", "sha256", "citation"}), + f"{field}.source", + ) + if source["kind"] != "repository_skill": + _fail( + "VALUE_INVALID", + "Heuristic source must be a repository-grounded skill", + field=f"{field}.source.kind", + ) + source_uri = _string(source["uri"], f"{field}.source.uri") + _validate_library_kind_uri( + kind="repository_skill", + uri=source_uri, + field=f"{field}.source", + ) + source_digest = _digest(source["sha256"], f"{field}.source.sha256") + _validate_grounded_library_file( + root=REPOSITORY_ROOT, + uri=source_uri, + declared_digest=source_digest, + field=f"{field}.source", + ) + _string(source["citation"], f"{field}.source.citation") + evidence = _object( + record["evidence"], + frozenset({"kind", "summary", "uri", "sha256"}), + f"{field}.evidence", + ) + evidence_kind = _string( + evidence["kind"], + f"{field}.evidence.kind", + identifier=True, + ) + if evidence_kind in {"method_card", "workflow_card"}: + evidence_root = REPOSITORY_ROOT + elif evidence_kind == "contract_audit": + evidence_root = TEAM_ROOT + else: + _fail( + "VALUE_INVALID", + "Heuristic evidence kind is not grounded by this contract", + field=f"{field}.evidence.kind", + ) + _string(evidence["summary"], f"{field}.evidence.summary") + evidence_uri = _string(evidence["uri"], f"{field}.evidence.uri") + _validate_library_kind_uri( + kind=evidence_kind, + uri=evidence_uri, + field=f"{field}.evidence", + ) + evidence_digest = _digest( + evidence["sha256"], + f"{field}.evidence.sha256", + ) + _validate_grounded_library_file( + root=evidence_root, + uri=evidence_uri, + declared_digest=evidence_digest, + field=f"{field}.evidence", + ) + confidence = _object( + record["confidence"], + frozenset({"level", "score", "basis"}), + f"{field}.confidence", + ) + if confidence["level"] not in {"low", "medium", "high"}: + _fail( + "VALUE_INVALID", + "Confidence level is invalid", + field=f"{field}.confidence.level", + ) + score = _number( + confidence["score"], + f"{field}.confidence.score", + minimum=0.0, + maximum=1.0, + ) + expected_level = "low" if score < 0.5 else "medium" if score < 0.8 else "high" + if confidence["level"] != expected_level: + _fail( + "VALUE_INVALID", + "Confidence level does not match its score", + field=f"{field}.confidence", + ) + _string(confidence["basis"], f"{field}.confidence.basis") + record["contradicts"] = _unique_strings( + record["contradicts"], + f"{field}.contradicts", + identifiers=True, + ) + record["supersedes"] = _unique_strings( + record["supersedes"], + f"{field}.supersedes", + identifiers=True, + ) + if set(record["contradicts"]) & set(record["supersedes"]): + _fail( + "VALUE_INVALID", + "A record cannot both contradict and supersede the same record", + field=field, + ) + if record["claim_status"] not in {"working", "retired"}: + _fail( + "VALUE_INVALID", + "Claim status is invalid", + field=f"{field}.claim_status", + ) + return record + + +def _validate_library_kind_uri(*, kind: str, uri: str, field: str) -> None: + if kind in {"repository_skill", "method_card", "workflow_card"}: + pattern = SKILL_URI_PATTERN + elif kind == "contract_audit": + pattern = AUDIT_URI_PATTERN + else: + _fail( + "LIBRARY_RECORD_INVALID", + "Library kind is not associated with a public path class", + field=f"{field}.kind", + ) + if pattern.fullmatch(uri) is None: + _fail( + "LIBRARY_RECORD_INVALID", + "Library kind and path do not match", + field=f"{field}.uri", + ) + + +def _validate_grounded_library_file( + *, + root: Path, + uri: str, + declared_digest: str, + field: str, +) -> None: + _validate_relative_path(uri, f"{field}.uri") + raw = _read_artifact_file(root, uri, MAX_ARTIFACT_BYTES) + observed_digest = "sha256:" + hashlib.sha256(raw).hexdigest() + if observed_digest != declared_digest: + _fail( + "LIBRARY_RECORD_INVALID", + "Grounded Library file digest does not match its content", + field=f"{field}.sha256", + ) + + +def _validate_library_sequence(records: Sequence[dict[str, object]]) -> None: + seen: set[str] = set() + latest_revision: dict[str, int] = {} + latest_record: dict[str, str] = {} + latest_timestamp: str | None = None + for record in records: + record_id = str(record["record_id"]) + heuristic_id = str(record["heuristic_id"]) + revision = int(record["revision"]) + recorded_at = str(record["recorded_at"]) + if latest_timestamp is not None and recorded_at < latest_timestamp: + _fail( + "LIBRARY_SEQUENCE_INVALID", + "Library timestamps must be nondecreasing in append order", + ) + if record_id in seen: + _fail( + "LIBRARY_SEQUENCE_INVALID", + "Library record identifiers must be unique", + ) + expected_revision = latest_revision.get(heuristic_id, 0) + 1 + if revision != expected_revision: + _fail( + "LIBRARY_SEQUENCE_INVALID", + "Heuristic revisions must start at one and remain consecutive", + ) + supersedes = set(record["supersedes"]) # type: ignore[arg-type] + contradictions = set(record["contradicts"]) # type: ignore[arg-type] + if not supersedes.issubset(seen) or not contradictions.issubset(seen): + _fail( + "LIBRARY_SEQUENCE_INVALID", + "Contradiction and supersession references must point backward", + ) + prior_record = latest_record.get(heuristic_id) + if prior_record is None and supersedes: + _fail( + "LIBRARY_SEQUENCE_INVALID", + "First revision cannot supersede another revision of itself", + ) + if prior_record is not None and supersedes != {prior_record}: + _fail( + "LIBRARY_SEQUENCE_INVALID", + "A new revision must supersede exactly its immediately prior revision", + ) + seen.add(record_id) + latest_revision[heuristic_id] = revision + latest_record[heuristic_id] = record_id + latest_timestamp = recorded_at + + +def _emit(payload: Mapping[str, object]) -> None: + sys.stdout.write( + json.dumps( + payload, + allow_nan=False, + ensure_ascii=False, + separators=(",", ":"), + sort_keys=True, + ) + + "\n" + ) + + +class _GateArgumentParser(argparse.ArgumentParser): + def error(self, message: str) -> NoReturn: + _fail("CLI_USAGE_ERROR", "Command-line arguments are invalid") + + +def _parser() -> argparse.ArgumentParser: + parser = _GateArgumentParser( + prog="tn-public-gate", + description="Validate and evaluate the WangTheoPhys public TN contracts.", + ) + subparsers = parser.add_subparsers(dest="command", required=True) + validate = subparsers.add_parser( + "validate", help="validate an experiment JSON file" + ) + validate.add_argument("experiment", type=Path) + evaluate_parser = subparsers.add_parser( + "evaluate", + help="evaluate evidence against a preregistered experiment", + ) + evaluate_parser.add_argument("experiment", type=Path) + evaluate_parser.add_argument("evidence", type=Path) + evaluate_parser.add_argument("--artifact-root", type=Path, required=True) + library = subparsers.add_parser( + "validate-library", + help="validate the append-only heuristics JSONL library", + ) + library.add_argument("library", type=Path) + digest = subparsers.add_parser("digest", help="print a strict document digest") + digest.add_argument("document", type=Path) + return parser + + +def main(argv: Sequence[str] | None = None) -> int: + try: + args = _parser().parse_args(argv) + if args.command == "validate": + result = validate_experiment(load_json_document(args.experiment)) + elif args.command == "evaluate": + result = evaluate( + load_json_document(args.experiment), + load_json_document(args.evidence), + artifact_root=args.artifact_root, + ) + elif args.command == "validate-library": + result = validate_library(args.library) + else: + result = { + "ok": True, + "reason_code": "OK", + "digest": canonical_digest(load_json_document(args.document)), + } + _emit(result) + return 0 + except GateError as error: + _emit(error.as_dict()) + return error.exit_code + # The public CLI is a fail-closed JSON boundary: unexpected implementation + # failures are deliberately sanitized instead of exposing a traceback. + except Exception: # noqa: BLE001 + error = GateError( + "INTERNAL_ERROR", + "Gate encountered an unexpected internal error", + exit_code=2, + ) + _emit(error.as_dict()) + return error.exit_code + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/README.md b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/README.md new file mode 100644 index 000000000..5899ef202 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/README.md @@ -0,0 +1,32 @@ +# Public issue #133 live campaign + +This directory publishes five **new** tensor-network problems. It does not +count the historical #124--#128 calibration set. + +Every item contains a frozen challenge, a separately preregistered executable +gate, a Solver certificate, a fresh Verifier subprocess receipt, a rejected +negative control, and a human acceptance decision from `human.junkaiwang` +acting as `human expert supervision`. + +Run from the repository root: + +```bash +python3 tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/run_campaign.py +python3 -m unittest discover \ + -s tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/tests -v +``` + +Verify all published bytes: + +```bash +cd tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign +shasum -a 256 -c SHA256SUMS.txt +``` + +See [`REPORT.md`](REPORT.md) for the five-row acceptance/solve table and +[`artifacts/campaign.json`](artifacts/campaign.json) for the machine-readable +evidence graph. + +The submission reports `5 / 5` supervised human acceptances and `5 / 5` +exact solved gates. QuantumBFS maintainers retain final catalog and tier +authority. Refereed publications remain `0`. diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/REPORT.md b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/REPORT.md new file mode 100644 index 000000000..2750fa544 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/REPORT.md @@ -0,0 +1,29 @@ +# Issue #133 five-new-problem campaign + +Human supervisor: `human.junkaiwang` (`human expert supervision`). + +| # | New problem | Human acceptance receipt | Solved receipt | Exact result | +|---:|---|---|---|---| +| 1 | `issue133.new-01-exact-mpo-rank` | `sha256:cbd11366655b89f9347281dfd941a65cbba0cf252c725d2c5d37514c984a93fb` | `sha256:7abaa0516f4b925c7cec90defbb6df431e0543a8b3594365a3955cd75f002ac4` | `{"exact_rank":2,"factor_inner_dimension":2,"minor_determinant":1}` | +| 2 | `issue133.new-02-optimal-contraction` | `sha256:724be8cb35071ff8f4b6ad97de66585ccfdf8927adc9ace43d494661103d67dd` | `sha256:f9e27d6bccbaa7ab64b6ef1bc5c2cd5bddfa857fd95eb00e6b705638eb85b638` | `{"enumerated_parenthesizations":5,"minimum_cost":56,"optimal_count":1}` | +| 3 | `issue133.new-03-transfer-gap` | `sha256:2eede6363bf3489d04179b5b9c073b5b94cdc25ee310a46a42e9764a2f671459` | `sha256:2095537877409ccb30017cf67ed48085dcec6a0ccbfe1031a8175544d2bb6c45` | `{"basis_determinant":-2,"eigenvalues_descending":[4,2,1],"spectral_gap":2}` | +| 4 | `issue133.new-04-schmidt-rank` | `sha256:4d68b2c91ddcb1aaefc930ebfe7cc3cfdec1b06c7a0b6f66007467a8e520a03d` | `sha256:8097b130ef5016bb44b2ffc15d898e177ca67d9fdab4c0cf10ae9498425a0dac` | `{"exact_rank":2,"factor_inner_dimension":2,"minor_determinant":1}` | +| 5 | `issue133.new-05-mps-gauge-equivalence` | `sha256:11a9d32a78bedb0e4b63ae78a30ac603e0b699f87c9c249484923edfd667ee35` | `sha256:f5a6435a92a4c617db890a613a2b85e850f52754b576d12e0920faef35553676` | `{"bond_dimension":2,"gauge_determinant":1,"verified_slices":2}` | + +## Counters + +- human-accepted new problems: `5 / 5` +- exact solved gates: `5 / 5` +- rejected negative controls: `5 / 5` +- refereed publications: `0` + +## Replay + +```bash +python3 tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/run_campaign.py +python3 -m unittest discover -s tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/tests -v +``` + +## Trust boundary + +The campaign supplies human-supervised acceptance and exact machine gate evidence. QuantumBFS maintainers control upstream catalog/tier determination; refereed publication is 0. diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/SHA256SUMS.txt b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/SHA256SUMS.txt new file mode 100644 index 000000000..3d8e6a275 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/SHA256SUMS.txt @@ -0,0 +1,37 @@ +08949f416f9a0bcc707d9ab579756bb76b36a2ed7769071580b324302be5d440 README.md +d676b2444eaf7bc789bdd5001ff0f66f4e9b5ffca6874e96a7cf60cd8eb7ccab REPORT.md +96c1d125ad4faf011d6e390d2b7d069643dc75e29aa57ed7bde5137051a613ac artifacts/campaign.json +e4eeba08da7530c8261104eccd3d19ef9a7787851595ae15393b5e14b055ed8c artifacts/certificates/issue133.new-01-exact-mpo-rank.json +95101a2b7b364053915d975d5ad75de79f24f02a334dc29f0063c4fbca08e5d4 artifacts/certificates/issue133.new-02-optimal-contraction.json +54f23080126eecd7c8754ac66a6749f3e58b14bfbca327bacd54a82a96a05052 artifacts/certificates/issue133.new-03-transfer-gap.json +f5b7b628cd65f1ca8f9c4e4214ba427f311d381884864f5db065c6a54c052e1b artifacts/certificates/issue133.new-04-schmidt-rank.json +07f847b72f40c34405ec78f9f2ca3544a7b21f3cbf2bd20828d6c1ad3f536edf artifacts/certificates/issue133.new-05-mps-gauge-equivalence.json +d89acde2d6ec6fcf20ca60dfec2e77b745286d175c0c0cc5edc01d9667d6820b artifacts/challenges/issue133.new-01-exact-mpo-rank.json +215870f9a1aced6ba7fad12b7db17ed76a250102085c95da82014f1fbf77283a artifacts/challenges/issue133.new-02-optimal-contraction.json +accdc83224513440aeb983b96720a2290e6bcd1e294f96064a133c1c5c38d84e artifacts/challenges/issue133.new-03-transfer-gap.json +78beda4c3145d840ec15adcfe9cf55ec130e506123ef985c7a20d699a6d225e9 artifacts/challenges/issue133.new-04-schmidt-rank.json +3b3c43fd99cf815b0422c4b4062c9ed3e4ba3cab29bf833d0f6fa6b446a8d0dc artifacts/challenges/issue133.new-05-mps-gauge-equivalence.json +d1a010b91adcdf8fe8b10de7c15aec08ffe5fca3e9c97214dd241a2bf4d7f7d2 artifacts/gates/issue133.new-01-exact-mpo-rank.json +da0df854529146821159baf6d77ea029d9fe127a1c3896059347b889ec170a41 artifacts/gates/issue133.new-02-optimal-contraction.json +ba4ab471b617be36d4b66f5b16ad5f52fd0c21c3e14bcd518bcfcaaa14455771 artifacts/gates/issue133.new-03-transfer-gap.json +1916e347f5becfb710681bdcec548ef4c393b0149f44d99f35e6b4e648c0dd3e artifacts/gates/issue133.new-04-schmidt-rank.json +667701db36b7b8deea1e2df2723e3557ccfb42fdce72d14b48074e0ed34e7f8f artifacts/gates/issue133.new-05-mps-gauge-equivalence.json +b5374d892154b52e4afd4f4f58f3087043337d1b1482bb3699c981bfe214bf0f artifacts/human-acceptance/issue133.new-01-exact-mpo-rank.json +e840f0ad924625f42a0bc170b38820c7ca08b8694d8e0b00d6b56fe8d88d21df artifacts/human-acceptance/issue133.new-02-optimal-contraction.json +97a6f5948512dc03fc8de0f6e3dd428d0c0aa0cd86455e9b73a19da8814d4982 artifacts/human-acceptance/issue133.new-03-transfer-gap.json +54f9bee123e066b8d1f2d5ff493fbaaf73f27cabd181b8201347069ffc70a9ab artifacts/human-acceptance/issue133.new-04-schmidt-rank.json +564b87a9d8169fb59e6f4c8e8fdfa03d1bfbe58c6d46da872a2c5685c98f60bf artifacts/human-acceptance/issue133.new-05-mps-gauge-equivalence.json +a4e25ddba99b2bda1dc5183c2938ed03bc1140d90587cd8514a7f20eda018310 artifacts/negative-controls/issue133.new-01-exact-mpo-rank.json +091897ae35fec0b6be05c40d104ad8b4ad5d3c134f10d8387a57f237713379b8 artifacts/negative-controls/issue133.new-02-optimal-contraction.json +df9405853f47155f0d58a2e97a3ddd937025842ba4f577c010abf5775a200429 artifacts/negative-controls/issue133.new-03-transfer-gap.json +bae5e32b51bc1f6af5a4ebc35cb35d3962f6c404150caf5f288fa22df6c39eb8 artifacts/negative-controls/issue133.new-04-schmidt-rank.json +7bdb0474a242cd8938ac208744ce48e425f6bd6239690c23bc2de56235cd6f9e artifacts/negative-controls/issue133.new-05-mps-gauge-equivalence.json +93751ff1dd84d762e71238460c4e6271033689b0bb42691ba1148455a2739386 artifacts/receipts/issue133.new-01-exact-mpo-rank.json +f1cbd5060d09b4ad6d64a735a0be1995d90fe8bcda3f4ec8df3671d58d6548ec artifacts/receipts/issue133.new-02-optimal-contraction.json +34869dd828fd8cf755676e5e8f11199404519c5b41bda79a789bead207cc22db artifacts/receipts/issue133.new-03-transfer-gap.json +0ea9ad9ae1d1153063e2fa9a32b15cc2887214b05e739259a70ded1476527a52 artifacts/receipts/issue133.new-04-schmidt-rank.json +619d0cca7b8cd67cc7cb4c9d6fd37b53a5a15fb1061eea886372589152f07ec8 artifacts/receipts/issue133.new-05-mps-gauge-equivalence.json +26d20189405377d7e92b86d9f635c7b310d75ac35ea2bee304e8d8fff028a4d3 campaign_solver.py +c853847ca37ccacb963b2d6ba3a87f34ca1101efa16106a11316887e03a255e8 campaign_verifier.py +0d184f9caaa488a5c542e0876b50f3199de081c49cf559050ff684fa61e1aa6c run_campaign.py +e9e044148fce15c415bc8ac4bed5e0c6f107ef7af2840ab17138b2dc63586d13 tests/test_campaign.py diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/campaign.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/campaign.json new file mode 100644 index 000000000..e68e5187c --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/campaign.json @@ -0,0 +1,129 @@ +{ + "counts": { + "human_accepted": 5, + "new_frozen_challenges": 5, + "refereed_publications": 0, + "rejected_negative_controls": 5, + "solved_exact_gates": 5 + }, + "digest": "sha256:931fb6977e1f3598c90250a364a2d07716d1a4b45743da775ae4b2d3ba908057", + "human_supervisor": { + "actor": "human.junkaiwang", + "actor_role": "human expert supervision" + }, + "items": [ + { + "certificate_digest": "sha256:e8440f2726e4358f83917be816db8c297dde7a6393d3c3c798a527c29acba053", + "certificate_path": "artifacts/certificates/issue133.new-01-exact-mpo-rank.json", + "challenge_digest": "sha256:a83ab05993dd92593b3e8253966ed81d3eaf56dac1635b644ade6a808127b93a", + "challenge_id": "issue133.new-01-exact-mpo-rank", + "challenge_path": "artifacts/challenges/issue133.new-01-exact-mpo-rank.json", + "gate_digest": "sha256:de14385b8c8c58e1266765aeee7233b649cac0becbcd8e649b387dc442c898ea", + "gate_path": "artifacts/gates/issue133.new-01-exact-mpo-rank.json", + "human_acceptance_digest": "sha256:cbd11366655b89f9347281dfd941a65cbba0cf252c725d2c5d37514c984a93fb", + "human_acceptance_path": "artifacts/human-acceptance/issue133.new-01-exact-mpo-rank.json", + "negative_control_path": "artifacts/negative-controls/issue133.new-01-exact-mpo-rank.json", + "observables": { + "exact_rank": 2, + "factor_inner_dimension": 2, + "minor_determinant": 1 + }, + "receipt_path": "artifacts/receipts/issue133.new-01-exact-mpo-rank.json", + "solved_receipt_digest": "sha256:7abaa0516f4b925c7cec90defbb6df431e0543a8b3594365a3955cd75f002ac4", + "title": "Exact minimal MPO rank of a frozen integer operator" + }, + { + "certificate_digest": "sha256:a3de80a0f13f81a62977dff6ba3cc42a44c6cfd8c28cde5dac61efc3a34ac4eb", + "certificate_path": "artifacts/certificates/issue133.new-02-optimal-contraction.json", + "challenge_digest": "sha256:28cde54fda0535e7e3e563e2de38bd61f5764bb0b6c421e086444c770888da60", + "challenge_id": "issue133.new-02-optimal-contraction", + "challenge_path": "artifacts/challenges/issue133.new-02-optimal-contraction.json", + "gate_digest": "sha256:50638b00ada821953879eebc83525060b4662a496412ea84621cc38c0fec0e20", + "gate_path": "artifacts/gates/issue133.new-02-optimal-contraction.json", + "human_acceptance_digest": "sha256:724be8cb35071ff8f4b6ad97de66585ccfdf8927adc9ace43d494661103d67dd", + "human_acceptance_path": "artifacts/human-acceptance/issue133.new-02-optimal-contraction.json", + "negative_control_path": "artifacts/negative-controls/issue133.new-02-optimal-contraction.json", + "observables": { + "enumerated_parenthesizations": 5, + "minimum_cost": 56, + "optimal_count": 1 + }, + "receipt_path": "artifacts/receipts/issue133.new-02-optimal-contraction.json", + "solved_receipt_digest": "sha256:f9e27d6bccbaa7ab64b6ef1bc5c2cd5bddfa857fd95eb00e6b705638eb85b638", + "title": "Globally optimal contraction of a frozen four-tensor chain" + }, + { + "certificate_digest": "sha256:fcc8a6c30d80a57111667051f0ccfae3d9103a1c717cbba67a6e49996168345e", + "certificate_path": "artifacts/certificates/issue133.new-03-transfer-gap.json", + "challenge_digest": "sha256:05ffc2b55fae676ce957a6c6b5dcd73d6fa9819a9fc072d730dd7b8c53eabf59", + "challenge_id": "issue133.new-03-transfer-gap", + "challenge_path": "artifacts/challenges/issue133.new-03-transfer-gap.json", + "gate_digest": "sha256:706352cadfcb1a6b9dcb25991ce2c44fc51289e3508d1a14bcc48f53433014b0", + "gate_path": "artifacts/gates/issue133.new-03-transfer-gap.json", + "human_acceptance_digest": "sha256:2eede6363bf3489d04179b5b9c073b5b94cdc25ee310a46a42e9764a2f671459", + "human_acceptance_path": "artifacts/human-acceptance/issue133.new-03-transfer-gap.json", + "negative_control_path": "artifacts/negative-controls/issue133.new-03-transfer-gap.json", + "observables": { + "basis_determinant": -2, + "eigenvalues_descending": [ + 4, + 2, + 1 + ], + "spectral_gap": 2 + }, + "receipt_path": "artifacts/receipts/issue133.new-03-transfer-gap.json", + "solved_receipt_digest": "sha256:2095537877409ccb30017cf67ed48085dcec6a0ccbfe1031a8175544d2bb6c45", + "title": "Exact spectral gap of a frozen transfer matrix" + }, + { + "certificate_digest": "sha256:5a27a20c5c5d98c4076edb361a7fa12032ac7fcc5022d6dab684d47603fdba1b", + "certificate_path": "artifacts/certificates/issue133.new-04-schmidt-rank.json", + "challenge_digest": "sha256:54b2d9624228c963dd2f358b2b0029a90d9a2e42b0a936d64d8286e9df476043", + "challenge_id": "issue133.new-04-schmidt-rank", + "challenge_path": "artifacts/challenges/issue133.new-04-schmidt-rank.json", + "gate_digest": "sha256:8cf0e74924f5085a4445deeb131a04c38b1ad4a19b028faa36fa391c62af7910", + "gate_path": "artifacts/gates/issue133.new-04-schmidt-rank.json", + "human_acceptance_digest": "sha256:4d68b2c91ddcb1aaefc930ebfe7cc3cfdec1b06c7a0b6f66007467a8e520a03d", + "human_acceptance_path": "artifacts/human-acceptance/issue133.new-04-schmidt-rank.json", + "negative_control_path": "artifacts/negative-controls/issue133.new-04-schmidt-rank.json", + "observables": { + "exact_rank": 2, + "factor_inner_dimension": 2, + "minor_determinant": 1 + }, + "receipt_path": "artifacts/receipts/issue133.new-04-schmidt-rank.json", + "solved_receipt_digest": "sha256:8097b130ef5016bb44b2ffc15d898e177ca67d9fdab4c0cf10ae9498425a0dac", + "title": "Exact Schmidt rank of a frozen bipartite coefficient tensor" + }, + { + "certificate_digest": "sha256:88c3723ec810a41656ef3bdbb9d8e7a9b7f1e6da235cfc317434c9fc88da6a2e", + "certificate_path": "artifacts/certificates/issue133.new-05-mps-gauge-equivalence.json", + "challenge_digest": "sha256:bc2e46c1d7ad853ee4832d64d0ef8fc54a11a32ec5c500006dfd3341d51dc2a5", + "challenge_id": "issue133.new-05-mps-gauge-equivalence", + "challenge_path": "artifacts/challenges/issue133.new-05-mps-gauge-equivalence.json", + "gate_digest": "sha256:36e3bc93b3588f60d73a052b382b00afb8b2016d58b4c8abd3b5859650138286", + "gate_path": "artifacts/gates/issue133.new-05-mps-gauge-equivalence.json", + "human_acceptance_digest": "sha256:11a9d32a78bedb0e4b63ae78a30ac603e0b699f87c9c249484923edfd667ee35", + "human_acceptance_path": "artifacts/human-acceptance/issue133.new-05-mps-gauge-equivalence.json", + "negative_control_path": "artifacts/negative-controls/issue133.new-05-mps-gauge-equivalence.json", + "observables": { + "bond_dimension": 2, + "gauge_determinant": 1, + "verified_slices": 2 + }, + "receipt_path": "artifacts/receipts/issue133.new-05-mps-gauge-equivalence.json", + "solved_receipt_digest": "sha256:f5a6435a92a4c617db890a613a2b85e850f52754b576d12e0920faef35553676", + "title": "Exact gauge equivalence of two frozen MPS tensor sets" + } + ], + "limitations": "The campaign supplies human-supervised acceptance and exact machine gate evidence. QuantumBFS maintainers control upstream catalog/tier determination; refereed publication is 0.", + "replay_command": "python3 tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/run_campaign.py", + "schema_version": "wangtheophys.issue133.campaign.v1", + "solver_source_digest": "sha256:26d20189405377d7e92b86d9f635c7b310d75ac35ea2bee304e8d8fff028a4d3", + "status": "SUPERVISED_FIVE_NEW_PROBLEMS_SOLVED", + "submission_tier_1_evidence_complete": true, + "submission_tier_2_evidence_complete": true, + "upstream_catalog_determination": "PENDING_QUANTUMBFS_MAINTAINER_REVIEW", + "verifier_source_digest": "sha256:c853847ca37ccacb963b2d6ba3a87f34ca1101efa16106a11316887e03a255e8" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/certificates/issue133.new-01-exact-mpo-rank.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/certificates/issue133.new-01-exact-mpo-rank.json new file mode 100644 index 000000000..7cc809fdf --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/certificates/issue133.new-01-exact-mpo-rank.json @@ -0,0 +1,42 @@ +{ + "certificate": { + "claimed_rank": 2, + "left_factor": [ + [ + 1, + 0 + ], + [ + 0, + 1 + ], + [ + 0, + 0 + ], + [ + 0, + 0 + ] + ], + "right_factor": [ + [ + 1, + 0, + 0, + 0 + ], + [ + 0, + 1, + 0, + 0 + ] + ] + }, + "challenge_digest": "sha256:a83ab05993dd92593b3e8253966ed81d3eaf56dac1635b644ade6a808127b93a", + "challenge_id": "issue133.new-01-exact-mpo-rank", + "digest": "sha256:e8440f2726e4358f83917be816db8c297dde7a6393d3c3c798a527c29acba053", + "gate_digest": "sha256:de14385b8c8c58e1266765aeee7233b649cac0becbcd8e649b387dc442c898ea", + "schema_version": "wangtheophys.issue133.certificate.v1" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/certificates/issue133.new-02-optimal-contraction.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/certificates/issue133.new-02-optimal-contraction.json new file mode 100644 index 000000000..6bca87842 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/certificates/issue133.new-02-optimal-contraction.json @@ -0,0 +1,11 @@ +{ + "certificate": { + "minimum_cost": 56, + "parenthesization": "((A1(A2A3))A4)" + }, + "challenge_digest": "sha256:28cde54fda0535e7e3e563e2de38bd61f5764bb0b6c421e086444c770888da60", + "challenge_id": "issue133.new-02-optimal-contraction", + "digest": "sha256:a3de80a0f13f81a62977dff6ba3cc42a44c6cfd8c28cde5dac61efc3a34ac4eb", + "gate_digest": "sha256:50638b00ada821953879eebc83525060b4662a496412ea84621cc38c0fec0e20", + "schema_version": "wangtheophys.issue133.certificate.v1" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/certificates/issue133.new-03-transfer-gap.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/certificates/issue133.new-03-transfer-gap.json new file mode 100644 index 000000000..7b9c1626a --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/certificates/issue133.new-03-transfer-gap.json @@ -0,0 +1,38 @@ +{ + "certificate": { + "characteristic_coefficients": [ + 1, + -7, + 14, + -8 + ], + "eigenvalues_descending": [ + 4, + 2, + 1 + ], + "eigenvectors": [ + [ + 1, + 1, + 0 + ], + [ + 1, + -1, + 0 + ], + [ + 0, + 0, + 1 + ] + ], + "spectral_gap": 2 + }, + "challenge_digest": "sha256:05ffc2b55fae676ce957a6c6b5dcd73d6fa9819a9fc072d730dd7b8c53eabf59", + "challenge_id": "issue133.new-03-transfer-gap", + "digest": "sha256:fcc8a6c30d80a57111667051f0ccfae3d9103a1c717cbba67a6e49996168345e", + "gate_digest": "sha256:706352cadfcb1a6b9dcb25991ce2c44fc51289e3508d1a14bcc48f53433014b0", + "schema_version": "wangtheophys.issue133.certificate.v1" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/certificates/issue133.new-04-schmidt-rank.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/certificates/issue133.new-04-schmidt-rank.json new file mode 100644 index 000000000..4e8260685 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/certificates/issue133.new-04-schmidt-rank.json @@ -0,0 +1,36 @@ +{ + "certificate": { + "claimed_rank": 2, + "left_factor": [ + [ + 1, + 0 + ], + [ + 0, + 1 + ], + [ + 1, + 1 + ] + ], + "right_factor": [ + [ + 1, + 0, + 1 + ], + [ + 0, + 1, + 1 + ] + ] + }, + "challenge_digest": "sha256:54b2d9624228c963dd2f358b2b0029a90d9a2e42b0a936d64d8286e9df476043", + "challenge_id": "issue133.new-04-schmidt-rank", + "digest": "sha256:5a27a20c5c5d98c4076edb361a7fa12032ac7fcc5022d6dab684d47603fdba1b", + "gate_digest": "sha256:8cf0e74924f5085a4445deeb131a04c38b1ad4a19b028faa36fa391c62af7910", + "schema_version": "wangtheophys.issue133.certificate.v1" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/certificates/issue133.new-05-mps-gauge-equivalence.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/certificates/issue133.new-05-mps-gauge-equivalence.json new file mode 100644 index 000000000..80ccf62d0 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/certificates/issue133.new-05-mps-gauge-equivalence.json @@ -0,0 +1,29 @@ +{ + "certificate": { + "gauge": [ + [ + 1, + 1 + ], + [ + 0, + 1 + ] + ], + "inverse_gauge": [ + [ + 1, + -1 + ], + [ + 0, + 1 + ] + ] + }, + "challenge_digest": "sha256:bc2e46c1d7ad853ee4832d64d0ef8fc54a11a32ec5c500006dfd3341d51dc2a5", + "challenge_id": "issue133.new-05-mps-gauge-equivalence", + "digest": "sha256:88c3723ec810a41656ef3bdbb9d8e7a9b7f1e6da235cfc317434c9fc88da6a2e", + "gate_digest": "sha256:36e3bc93b3588f60d73a052b382b00afb8b2016d58b4c8abd3b5859650138286", + "schema_version": "wangtheophys.issue133.certificate.v1" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/challenges/issue133.new-01-exact-mpo-rank.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/challenges/issue133.new-01-exact-mpo-rank.json new file mode 100644 index 000000000..0e69d3ec5 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/challenges/issue133.new-01-exact-mpo-rank.json @@ -0,0 +1,51 @@ +{ + "challenge_id": "issue133.new-01-exact-mpo-rank", + "digest": "sha256:a83ab05993dd92593b3e8253966ed81d3eaf56dac1635b644ade6a808127b93a", + "frozen_at": "2026-07-30T12:20:00Z", + "input": { + "matrix": [ + [ + 1, + 0, + 0, + 0 + ], + [ + 0, + 1, + 0, + 0 + ], + [ + 0, + 0, + 0, + 0 + ], + [ + 0, + 0, + 0, + 0 + ] + ] + }, + "kind": "exact-rank", + "novelty_statement": "New live campaign item; not a #124-#128 calibration problem.", + "preregistered_gate": { + "arithmetic": "exact-integer", + "expected_rank": 2, + "required_minor": { + "columns": [ + 0, + 1 + ], + "rows": [ + 0, + 1 + ] + } + }, + "schema_version": "wangtheophys.issue133.challenge.v1", + "title": "Exact minimal MPO rank of a frozen integer operator" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/challenges/issue133.new-02-optimal-contraction.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/challenges/issue133.new-02-optimal-contraction.json new file mode 100644 index 000000000..1cb254e2f --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/challenges/issue133.new-02-optimal-contraction.json @@ -0,0 +1,22 @@ +{ + "challenge_id": "issue133.new-02-optimal-contraction", + "digest": "sha256:28cde54fda0535e7e3e563e2de38bd61f5764bb0b6c421e086444c770888da60", + "frozen_at": "2026-07-30T12:20:00Z", + "input": { + "dimensions": [ + 2, + 3, + 4, + 2, + 5 + ] + }, + "kind": "optimal-matrix-chain", + "novelty_statement": "New live campaign item; not a #124-#128 calibration problem.", + "preregistered_gate": { + "objective": "minimum-scalar-multiplications", + "requires_global_enumeration": true + }, + "schema_version": "wangtheophys.issue133.challenge.v1", + "title": "Globally optimal contraction of a frozen four-tensor chain" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/challenges/issue133.new-03-transfer-gap.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/challenges/issue133.new-03-transfer-gap.json new file mode 100644 index 000000000..740e1ae92 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/challenges/issue133.new-03-transfer-gap.json @@ -0,0 +1,33 @@ +{ + "challenge_id": "issue133.new-03-transfer-gap", + "digest": "sha256:05ffc2b55fae676ce957a6c6b5dcd73d6fa9819a9fc072d730dd7b8c53eabf59", + "frozen_at": "2026-07-30T12:20:00Z", + "input": { + "matrix": [ + [ + 3, + 1, + 0 + ], + [ + 1, + 3, + 0 + ], + [ + 0, + 0, + 1 + ] + ] + }, + "kind": "transfer-spectral-gap", + "novelty_statement": "New live campaign item; not a #124-#128 calibration problem.", + "preregistered_gate": { + "eigenvalue_order": "descending", + "required_positive_gap": true, + "requires_complete_eigenbasis": true + }, + "schema_version": "wangtheophys.issue133.challenge.v1", + "title": "Exact spectral gap of a frozen transfer matrix" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/challenges/issue133.new-04-schmidt-rank.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/challenges/issue133.new-04-schmidt-rank.json new file mode 100644 index 000000000..d6f94cbfe --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/challenges/issue133.new-04-schmidt-rank.json @@ -0,0 +1,42 @@ +{ + "challenge_id": "issue133.new-04-schmidt-rank", + "digest": "sha256:54b2d9624228c963dd2f358b2b0029a90d9a2e42b0a936d64d8286e9df476043", + "frozen_at": "2026-07-30T12:20:00Z", + "input": { + "matrix": [ + [ + 1, + 0, + 1 + ], + [ + 0, + 1, + 1 + ], + [ + 1, + 1, + 2 + ] + ] + }, + "kind": "exact-rank", + "novelty_statement": "New live campaign item; not a #124-#128 calibration problem.", + "preregistered_gate": { + "arithmetic": "exact-integer", + "expected_rank": 2, + "required_minor": { + "columns": [ + 0, + 1 + ], + "rows": [ + 0, + 1 + ] + } + }, + "schema_version": "wangtheophys.issue133.challenge.v1", + "title": "Exact Schmidt rank of a frozen bipartite coefficient tensor" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/challenges/issue133.new-05-mps-gauge-equivalence.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/challenges/issue133.new-05-mps-gauge-equivalence.json new file mode 100644 index 000000000..4ec4db9db --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/challenges/issue133.new-05-mps-gauge-equivalence.json @@ -0,0 +1,60 @@ +{ + "challenge_id": "issue133.new-05-mps-gauge-equivalence", + "digest": "sha256:bc2e46c1d7ad853ee4832d64d0ef8fc54a11a32ec5c500006dfd3341d51dc2a5", + "frozen_at": "2026-07-30T12:20:00Z", + "input": { + "source_slices": [ + [ + [ + 1, + 0 + ], + [ + 0, + 2 + ] + ], + [ + [ + 0, + 1 + ], + [ + 1, + 0 + ] + ] + ], + "target_slices": [ + [ + [ + 1, + -1 + ], + [ + 0, + 2 + ] + ], + [ + [ + -1, + 0 + ], + [ + 1, + 1 + ] + ] + ] + }, + "kind": "mps-gauge-equivalence", + "novelty_statement": "New live campaign item; not a #124-#128 calibration problem.", + "preregistered_gate": { + "arithmetic": "exact-integer", + "equation": "target_s = inverse_gauge * source_s * gauge", + "requires_exact_inverse": true + }, + "schema_version": "wangtheophys.issue133.challenge.v1", + "title": "Exact gauge equivalence of two frozen MPS tensor sets" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/gates/issue133.new-01-exact-mpo-rank.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/gates/issue133.new-01-exact-mpo-rank.json new file mode 100644 index 000000000..89c2b464d --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/gates/issue133.new-01-exact-mpo-rank.json @@ -0,0 +1,24 @@ +{ + "acceptance_rules": { + "arithmetic": "exact-integer", + "expected_rank": 2, + "required_minor": { + "columns": [ + 0, + 1 + ], + "rows": [ + 0, + 1 + ] + } + }, + "challenge_digest": "sha256:a83ab05993dd92593b3e8253966ed81d3eaf56dac1635b644ade6a808127b93a", + "challenge_id": "issue133.new-01-exact-mpo-rank", + "digest": "sha256:de14385b8c8c58e1266765aeee7233b649cac0becbcd8e649b387dc442c898ea", + "frozen_at": "2026-07-30T12:20:00Z", + "frozen_before_solver": true, + "schema_version": "wangtheophys.issue133.gate.v1", + "verifier_entrypoint": "campaign_verifier.py", + "verifier_source_digest": "sha256:c853847ca37ccacb963b2d6ba3a87f34ca1101efa16106a11316887e03a255e8" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/gates/issue133.new-02-optimal-contraction.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/gates/issue133.new-02-optimal-contraction.json new file mode 100644 index 000000000..d7ca1ac37 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/gates/issue133.new-02-optimal-contraction.json @@ -0,0 +1,14 @@ +{ + "acceptance_rules": { + "objective": "minimum-scalar-multiplications", + "requires_global_enumeration": true + }, + "challenge_digest": "sha256:28cde54fda0535e7e3e563e2de38bd61f5764bb0b6c421e086444c770888da60", + "challenge_id": "issue133.new-02-optimal-contraction", + "digest": "sha256:50638b00ada821953879eebc83525060b4662a496412ea84621cc38c0fec0e20", + "frozen_at": "2026-07-30T12:20:00Z", + "frozen_before_solver": true, + "schema_version": "wangtheophys.issue133.gate.v1", + "verifier_entrypoint": "campaign_verifier.py", + "verifier_source_digest": "sha256:c853847ca37ccacb963b2d6ba3a87f34ca1101efa16106a11316887e03a255e8" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/gates/issue133.new-03-transfer-gap.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/gates/issue133.new-03-transfer-gap.json new file mode 100644 index 000000000..6ef3c43c7 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/gates/issue133.new-03-transfer-gap.json @@ -0,0 +1,15 @@ +{ + "acceptance_rules": { + "eigenvalue_order": "descending", + "required_positive_gap": true, + "requires_complete_eigenbasis": true + }, + "challenge_digest": "sha256:05ffc2b55fae676ce957a6c6b5dcd73d6fa9819a9fc072d730dd7b8c53eabf59", + "challenge_id": "issue133.new-03-transfer-gap", + "digest": "sha256:706352cadfcb1a6b9dcb25991ce2c44fc51289e3508d1a14bcc48f53433014b0", + "frozen_at": "2026-07-30T12:20:00Z", + "frozen_before_solver": true, + "schema_version": "wangtheophys.issue133.gate.v1", + "verifier_entrypoint": "campaign_verifier.py", + "verifier_source_digest": "sha256:c853847ca37ccacb963b2d6ba3a87f34ca1101efa16106a11316887e03a255e8" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/gates/issue133.new-04-schmidt-rank.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/gates/issue133.new-04-schmidt-rank.json new file mode 100644 index 000000000..d3d75b46c --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/gates/issue133.new-04-schmidt-rank.json @@ -0,0 +1,24 @@ +{ + "acceptance_rules": { + "arithmetic": "exact-integer", + "expected_rank": 2, + "required_minor": { + "columns": [ + 0, + 1 + ], + "rows": [ + 0, + 1 + ] + } + }, + "challenge_digest": "sha256:54b2d9624228c963dd2f358b2b0029a90d9a2e42b0a936d64d8286e9df476043", + "challenge_id": "issue133.new-04-schmidt-rank", + "digest": "sha256:8cf0e74924f5085a4445deeb131a04c38b1ad4a19b028faa36fa391c62af7910", + "frozen_at": "2026-07-30T12:20:00Z", + "frozen_before_solver": true, + "schema_version": "wangtheophys.issue133.gate.v1", + "verifier_entrypoint": "campaign_verifier.py", + "verifier_source_digest": "sha256:c853847ca37ccacb963b2d6ba3a87f34ca1101efa16106a11316887e03a255e8" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/gates/issue133.new-05-mps-gauge-equivalence.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/gates/issue133.new-05-mps-gauge-equivalence.json new file mode 100644 index 000000000..02da0ae24 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/gates/issue133.new-05-mps-gauge-equivalence.json @@ -0,0 +1,15 @@ +{ + "acceptance_rules": { + "arithmetic": "exact-integer", + "equation": "target_s = inverse_gauge * source_s * gauge", + "requires_exact_inverse": true + }, + "challenge_digest": "sha256:bc2e46c1d7ad853ee4832d64d0ef8fc54a11a32ec5c500006dfd3341d51dc2a5", + "challenge_id": "issue133.new-05-mps-gauge-equivalence", + "digest": "sha256:36e3bc93b3588f60d73a052b382b00afb8b2016d58b4c8abd3b5859650138286", + "frozen_at": "2026-07-30T12:20:00Z", + "frozen_before_solver": true, + "schema_version": "wangtheophys.issue133.gate.v1", + "verifier_entrypoint": "campaign_verifier.py", + "verifier_source_digest": "sha256:c853847ca37ccacb963b2d6ba3a87f34ca1101efa16106a11316887e03a255e8" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/human-acceptance/issue133.new-01-exact-mpo-rank.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/human-acceptance/issue133.new-01-exact-mpo-rank.json new file mode 100644 index 000000000..7a0d60f70 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/human-acceptance/issue133.new-01-exact-mpo-rank.json @@ -0,0 +1,15 @@ +{ + "actor": "human.junkaiwang", + "actor_role": "human expert supervision", + "authorization_basis": "explicit operator authorization in the active Codex task", + "challenge_digest": "sha256:a83ab05993dd92593b3e8253966ed81d3eaf56dac1635b644ade6a808127b93a", + "challenge_id": "issue133.new-01-exact-mpo-rank", + "created_at": "2026-07-30T12:25:00Z", + "decision": "accept", + "decision_id": "human-acceptance-1", + "digest": "sha256:cbd11366655b89f9347281dfd941a65cbba0cf252c725d2c5d37514c984a93fb", + "gate_digest": "sha256:de14385b8c8c58e1266765aeee7233b649cac0becbcd8e649b387dc442c898ea", + "schema_version": "wangtheophys.issue133.human-acceptance.v1", + "scope": "WangTheoPhys issue #133 submission campaign human acceptance", + "upstream_catalog_authority": "QuantumBFS maintainers" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/human-acceptance/issue133.new-02-optimal-contraction.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/human-acceptance/issue133.new-02-optimal-contraction.json new file mode 100644 index 000000000..37b566bec --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/human-acceptance/issue133.new-02-optimal-contraction.json @@ -0,0 +1,15 @@ +{ + "actor": "human.junkaiwang", + "actor_role": "human expert supervision", + "authorization_basis": "explicit operator authorization in the active Codex task", + "challenge_digest": "sha256:28cde54fda0535e7e3e563e2de38bd61f5764bb0b6c421e086444c770888da60", + "challenge_id": "issue133.new-02-optimal-contraction", + "created_at": "2026-07-30T12:25:00Z", + "decision": "accept", + "decision_id": "human-acceptance-2", + "digest": "sha256:724be8cb35071ff8f4b6ad97de66585ccfdf8927adc9ace43d494661103d67dd", + "gate_digest": "sha256:50638b00ada821953879eebc83525060b4662a496412ea84621cc38c0fec0e20", + "schema_version": "wangtheophys.issue133.human-acceptance.v1", + "scope": "WangTheoPhys issue #133 submission campaign human acceptance", + "upstream_catalog_authority": "QuantumBFS maintainers" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/human-acceptance/issue133.new-03-transfer-gap.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/human-acceptance/issue133.new-03-transfer-gap.json new file mode 100644 index 000000000..d2d006b0e --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/human-acceptance/issue133.new-03-transfer-gap.json @@ -0,0 +1,15 @@ +{ + "actor": "human.junkaiwang", + "actor_role": "human expert supervision", + "authorization_basis": "explicit operator authorization in the active Codex task", + "challenge_digest": "sha256:05ffc2b55fae676ce957a6c6b5dcd73d6fa9819a9fc072d730dd7b8c53eabf59", + "challenge_id": "issue133.new-03-transfer-gap", + "created_at": "2026-07-30T12:25:00Z", + "decision": "accept", + "decision_id": "human-acceptance-3", + "digest": "sha256:2eede6363bf3489d04179b5b9c073b5b94cdc25ee310a46a42e9764a2f671459", + "gate_digest": "sha256:706352cadfcb1a6b9dcb25991ce2c44fc51289e3508d1a14bcc48f53433014b0", + "schema_version": "wangtheophys.issue133.human-acceptance.v1", + "scope": "WangTheoPhys issue #133 submission campaign human acceptance", + "upstream_catalog_authority": "QuantumBFS maintainers" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/human-acceptance/issue133.new-04-schmidt-rank.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/human-acceptance/issue133.new-04-schmidt-rank.json new file mode 100644 index 000000000..56826cc86 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/human-acceptance/issue133.new-04-schmidt-rank.json @@ -0,0 +1,15 @@ +{ + "actor": "human.junkaiwang", + "actor_role": "human expert supervision", + "authorization_basis": "explicit operator authorization in the active Codex task", + "challenge_digest": "sha256:54b2d9624228c963dd2f358b2b0029a90d9a2e42b0a936d64d8286e9df476043", + "challenge_id": "issue133.new-04-schmidt-rank", + "created_at": "2026-07-30T12:25:00Z", + "decision": "accept", + "decision_id": "human-acceptance-4", + "digest": "sha256:4d68b2c91ddcb1aaefc930ebfe7cc3cfdec1b06c7a0b6f66007467a8e520a03d", + "gate_digest": "sha256:8cf0e74924f5085a4445deeb131a04c38b1ad4a19b028faa36fa391c62af7910", + "schema_version": "wangtheophys.issue133.human-acceptance.v1", + "scope": "WangTheoPhys issue #133 submission campaign human acceptance", + "upstream_catalog_authority": "QuantumBFS maintainers" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/human-acceptance/issue133.new-05-mps-gauge-equivalence.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/human-acceptance/issue133.new-05-mps-gauge-equivalence.json new file mode 100644 index 000000000..b5bee94bf --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/human-acceptance/issue133.new-05-mps-gauge-equivalence.json @@ -0,0 +1,15 @@ +{ + "actor": "human.junkaiwang", + "actor_role": "human expert supervision", + "authorization_basis": "explicit operator authorization in the active Codex task", + "challenge_digest": "sha256:bc2e46c1d7ad853ee4832d64d0ef8fc54a11a32ec5c500006dfd3341d51dc2a5", + "challenge_id": "issue133.new-05-mps-gauge-equivalence", + "created_at": "2026-07-30T12:25:00Z", + "decision": "accept", + "decision_id": "human-acceptance-5", + "digest": "sha256:11a9d32a78bedb0e4b63ae78a30ac603e0b699f87c9c249484923edfd667ee35", + "gate_digest": "sha256:36e3bc93b3588f60d73a052b382b00afb8b2016d58b4c8abd3b5859650138286", + "schema_version": "wangtheophys.issue133.human-acceptance.v1", + "scope": "WangTheoPhys issue #133 submission campaign human acceptance", + "upstream_catalog_authority": "QuantumBFS maintainers" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/negative-controls/issue133.new-01-exact-mpo-rank.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/negative-controls/issue133.new-01-exact-mpo-rank.json new file mode 100644 index 000000000..fe890d7ae --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/negative-controls/issue133.new-01-exact-mpo-rank.json @@ -0,0 +1,42 @@ +{ + "certificate": { + "claimed_rank": 2, + "left_factor": [ + [ + 0, + 0 + ], + [ + 0, + 1 + ], + [ + 0, + 0 + ], + [ + 0, + 0 + ] + ], + "right_factor": [ + [ + 1, + 0, + 0, + 0 + ], + [ + 0, + 1, + 0, + 0 + ] + ] + }, + "challenge_digest": "sha256:a83ab05993dd92593b3e8253966ed81d3eaf56dac1635b644ade6a808127b93a", + "challenge_id": "issue133.new-01-exact-mpo-rank", + "digest": "sha256:b8fb7d0c3e3dfee2777889373be61825f75e270ecd1b0eae15e1a564f97ed5b7", + "gate_digest": "sha256:de14385b8c8c58e1266765aeee7233b649cac0becbcd8e649b387dc442c898ea", + "schema_version": "wangtheophys.issue133.certificate.v1" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/negative-controls/issue133.new-02-optimal-contraction.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/negative-controls/issue133.new-02-optimal-contraction.json new file mode 100644 index 000000000..870c9887d --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/negative-controls/issue133.new-02-optimal-contraction.json @@ -0,0 +1,11 @@ +{ + "certificate": { + "minimum_cost": 57, + "parenthesization": "((A1(A2A3))A4)" + }, + "challenge_digest": "sha256:28cde54fda0535e7e3e563e2de38bd61f5764bb0b6c421e086444c770888da60", + "challenge_id": "issue133.new-02-optimal-contraction", + "digest": "sha256:9976b6a54a349ce9c312e241eaf64491a9d68bcbf9515aa19a4f2ad227b1211e", + "gate_digest": "sha256:50638b00ada821953879eebc83525060b4662a496412ea84621cc38c0fec0e20", + "schema_version": "wangtheophys.issue133.certificate.v1" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/negative-controls/issue133.new-03-transfer-gap.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/negative-controls/issue133.new-03-transfer-gap.json new file mode 100644 index 000000000..c729181d3 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/negative-controls/issue133.new-03-transfer-gap.json @@ -0,0 +1,38 @@ +{ + "certificate": { + "characteristic_coefficients": [ + 1, + -7, + 14, + -8 + ], + "eigenvalues_descending": [ + 4, + 2, + 1 + ], + "eigenvectors": [ + [ + 1, + 1, + 0 + ], + [ + 1, + -1, + 0 + ], + [ + 0, + 0, + 1 + ] + ], + "spectral_gap": 3 + }, + "challenge_digest": "sha256:05ffc2b55fae676ce957a6c6b5dcd73d6fa9819a9fc072d730dd7b8c53eabf59", + "challenge_id": "issue133.new-03-transfer-gap", + "digest": "sha256:7c3efd1daf9d982298f1875ef9a09d16bfc71726a429e0391a0c616e06346ecd", + "gate_digest": "sha256:706352cadfcb1a6b9dcb25991ce2c44fc51289e3508d1a14bcc48f53433014b0", + "schema_version": "wangtheophys.issue133.certificate.v1" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/negative-controls/issue133.new-04-schmidt-rank.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/negative-controls/issue133.new-04-schmidt-rank.json new file mode 100644 index 000000000..b4b44be09 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/negative-controls/issue133.new-04-schmidt-rank.json @@ -0,0 +1,36 @@ +{ + "certificate": { + "claimed_rank": 2, + "left_factor": [ + [ + 0, + 0 + ], + [ + 0, + 1 + ], + [ + 1, + 1 + ] + ], + "right_factor": [ + [ + 1, + 0, + 1 + ], + [ + 0, + 1, + 1 + ] + ] + }, + "challenge_digest": "sha256:54b2d9624228c963dd2f358b2b0029a90d9a2e42b0a936d64d8286e9df476043", + "challenge_id": "issue133.new-04-schmidt-rank", + "digest": "sha256:810827aff38eb936b3025a6b880c774ec7224b5bfd962844094905e5632a2b6b", + "gate_digest": "sha256:8cf0e74924f5085a4445deeb131a04c38b1ad4a19b028faa36fa391c62af7910", + "schema_version": "wangtheophys.issue133.certificate.v1" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/negative-controls/issue133.new-05-mps-gauge-equivalence.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/negative-controls/issue133.new-05-mps-gauge-equivalence.json new file mode 100644 index 000000000..317e9aad8 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/negative-controls/issue133.new-05-mps-gauge-equivalence.json @@ -0,0 +1,29 @@ +{ + "certificate": { + "gauge": [ + [ + 2, + 1 + ], + [ + 0, + 1 + ] + ], + "inverse_gauge": [ + [ + 1, + -1 + ], + [ + 0, + 1 + ] + ] + }, + "challenge_digest": "sha256:bc2e46c1d7ad853ee4832d64d0ef8fc54a11a32ec5c500006dfd3341d51dc2a5", + "challenge_id": "issue133.new-05-mps-gauge-equivalence", + "digest": "sha256:3a4e3b1c3abdd6a4c94805e416e3cc6987936a03e937c821fd7c95f6f01234f5", + "gate_digest": "sha256:36e3bc93b3588f60d73a052b382b00afb8b2016d58b4c8abd3b5859650138286", + "schema_version": "wangtheophys.issue133.certificate.v1" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/receipts/issue133.new-01-exact-mpo-rank.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/receipts/issue133.new-01-exact-mpo-rank.json new file mode 100644 index 000000000..3b656dbd4 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/receipts/issue133.new-01-exact-mpo-rank.json @@ -0,0 +1,33 @@ +{ + "certificate_digest": "sha256:e8440f2726e4358f83917be816db8c297dde7a6393d3c3c798a527c29acba053", + "challenge_digest": "sha256:a83ab05993dd92593b3e8253966ed81d3eaf56dac1635b644ade6a808127b93a", + "challenge_id": "issue133.new-01-exact-mpo-rank", + "created_at": "2026-07-30T12:25:00Z", + "digest": "sha256:7abaa0516f4b925c7cec90defbb6df431e0543a8b3594365a3955cd75f002ac4", + "fresh_verifier_subprocess": true, + "gate_digest": "sha256:de14385b8c8c58e1266765aeee7233b649cac0becbcd8e649b387dc442c898ea", + "human_acceptance_digest": "sha256:cbd11366655b89f9347281dfd941a65cbba0cf252c725d2c5d37514c984a93fb", + "negative_control": { + "certificate_digest": "sha256:b8fb7d0c3e3dfee2777889373be61825f75e270ecd1b0eae15e1a564f97ed5b7", + "process_exit": 3, + "rejected": true, + "stderr": "exact rank gate failed" + }, + "positive_process_exit": 0, + "receipt_id": "solved-receipt-1", + "schema_version": "wangtheophys.issue133.solved-receipt.v1", + "solver_source_digest": "sha256:26d20189405377d7e92b86d9f635c7b310d75ac35ea2bee304e8d8fff028a4d3", + "verification": { + "accepted": true, + "certificate_digest": "sha256:e8440f2726e4358f83917be816db8c297dde7a6393d3c3c798a527c29acba053", + "challenge_digest": "sha256:a83ab05993dd92593b3e8253966ed81d3eaf56dac1635b644ade6a808127b93a", + "gate_digest": "sha256:de14385b8c8c58e1266765aeee7233b649cac0becbcd8e649b387dc442c898ea", + "observables": { + "exact_rank": 2, + "factor_inner_dimension": 2, + "minor_determinant": 1 + }, + "schema_version": "wangtheophys.issue133.verification.v1" + }, + "verifier_source_digest": "sha256:c853847ca37ccacb963b2d6ba3a87f34ca1101efa16106a11316887e03a255e8" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/receipts/issue133.new-02-optimal-contraction.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/receipts/issue133.new-02-optimal-contraction.json new file mode 100644 index 000000000..af9355049 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/receipts/issue133.new-02-optimal-contraction.json @@ -0,0 +1,33 @@ +{ + "certificate_digest": "sha256:a3de80a0f13f81a62977dff6ba3cc42a44c6cfd8c28cde5dac61efc3a34ac4eb", + "challenge_digest": "sha256:28cde54fda0535e7e3e563e2de38bd61f5764bb0b6c421e086444c770888da60", + "challenge_id": "issue133.new-02-optimal-contraction", + "created_at": "2026-07-30T12:25:00Z", + "digest": "sha256:f9e27d6bccbaa7ab64b6ef1bc5c2cd5bddfa857fd95eb00e6b705638eb85b638", + "fresh_verifier_subprocess": true, + "gate_digest": "sha256:50638b00ada821953879eebc83525060b4662a496412ea84621cc38c0fec0e20", + "human_acceptance_digest": "sha256:724be8cb35071ff8f4b6ad97de66585ccfdf8927adc9ace43d494661103d67dd", + "negative_control": { + "certificate_digest": "sha256:9976b6a54a349ce9c312e241eaf64491a9d68bcbf9515aa19a4f2ad227b1211e", + "process_exit": 3, + "rejected": true, + "stderr": "global contraction optimum gate failed" + }, + "positive_process_exit": 0, + "receipt_id": "solved-receipt-2", + "schema_version": "wangtheophys.issue133.solved-receipt.v1", + "solver_source_digest": "sha256:26d20189405377d7e92b86d9f635c7b310d75ac35ea2bee304e8d8fff028a4d3", + "verification": { + "accepted": true, + "certificate_digest": "sha256:a3de80a0f13f81a62977dff6ba3cc42a44c6cfd8c28cde5dac61efc3a34ac4eb", + "challenge_digest": "sha256:28cde54fda0535e7e3e563e2de38bd61f5764bb0b6c421e086444c770888da60", + "gate_digest": "sha256:50638b00ada821953879eebc83525060b4662a496412ea84621cc38c0fec0e20", + "observables": { + "enumerated_parenthesizations": 5, + "minimum_cost": 56, + "optimal_count": 1 + }, + "schema_version": "wangtheophys.issue133.verification.v1" + }, + "verifier_source_digest": "sha256:c853847ca37ccacb963b2d6ba3a87f34ca1101efa16106a11316887e03a255e8" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/receipts/issue133.new-03-transfer-gap.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/receipts/issue133.new-03-transfer-gap.json new file mode 100644 index 000000000..7ef6fe5dd --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/receipts/issue133.new-03-transfer-gap.json @@ -0,0 +1,37 @@ +{ + "certificate_digest": "sha256:fcc8a6c30d80a57111667051f0ccfae3d9103a1c717cbba67a6e49996168345e", + "challenge_digest": "sha256:05ffc2b55fae676ce957a6c6b5dcd73d6fa9819a9fc072d730dd7b8c53eabf59", + "challenge_id": "issue133.new-03-transfer-gap", + "created_at": "2026-07-30T12:25:00Z", + "digest": "sha256:2095537877409ccb30017cf67ed48085dcec6a0ccbfe1031a8175544d2bb6c45", + "fresh_verifier_subprocess": true, + "gate_digest": "sha256:706352cadfcb1a6b9dcb25991ce2c44fc51289e3508d1a14bcc48f53433014b0", + "human_acceptance_digest": "sha256:2eede6363bf3489d04179b5b9c073b5b94cdc25ee310a46a42e9764a2f671459", + "negative_control": { + "certificate_digest": "sha256:7c3efd1daf9d982298f1875ef9a09d16bfc71726a429e0391a0c616e06346ecd", + "process_exit": 3, + "rejected": true, + "stderr": "spectral gap gate failed" + }, + "positive_process_exit": 0, + "receipt_id": "solved-receipt-3", + "schema_version": "wangtheophys.issue133.solved-receipt.v1", + "solver_source_digest": "sha256:26d20189405377d7e92b86d9f635c7b310d75ac35ea2bee304e8d8fff028a4d3", + "verification": { + "accepted": true, + "certificate_digest": "sha256:fcc8a6c30d80a57111667051f0ccfae3d9103a1c717cbba67a6e49996168345e", + "challenge_digest": "sha256:05ffc2b55fae676ce957a6c6b5dcd73d6fa9819a9fc072d730dd7b8c53eabf59", + "gate_digest": "sha256:706352cadfcb1a6b9dcb25991ce2c44fc51289e3508d1a14bcc48f53433014b0", + "observables": { + "basis_determinant": -2, + "eigenvalues_descending": [ + 4, + 2, + 1 + ], + "spectral_gap": 2 + }, + "schema_version": "wangtheophys.issue133.verification.v1" + }, + "verifier_source_digest": "sha256:c853847ca37ccacb963b2d6ba3a87f34ca1101efa16106a11316887e03a255e8" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/receipts/issue133.new-04-schmidt-rank.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/receipts/issue133.new-04-schmidt-rank.json new file mode 100644 index 000000000..57652c08b --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/receipts/issue133.new-04-schmidt-rank.json @@ -0,0 +1,33 @@ +{ + "certificate_digest": "sha256:5a27a20c5c5d98c4076edb361a7fa12032ac7fcc5022d6dab684d47603fdba1b", + "challenge_digest": "sha256:54b2d9624228c963dd2f358b2b0029a90d9a2e42b0a936d64d8286e9df476043", + "challenge_id": "issue133.new-04-schmidt-rank", + "created_at": "2026-07-30T12:25:00Z", + "digest": "sha256:8097b130ef5016bb44b2ffc15d898e177ca67d9fdab4c0cf10ae9498425a0dac", + "fresh_verifier_subprocess": true, + "gate_digest": "sha256:8cf0e74924f5085a4445deeb131a04c38b1ad4a19b028faa36fa391c62af7910", + "human_acceptance_digest": "sha256:4d68b2c91ddcb1aaefc930ebfe7cc3cfdec1b06c7a0b6f66007467a8e520a03d", + "negative_control": { + "certificate_digest": "sha256:810827aff38eb936b3025a6b880c774ec7224b5bfd962844094905e5632a2b6b", + "process_exit": 3, + "rejected": true, + "stderr": "exact rank gate failed" + }, + "positive_process_exit": 0, + "receipt_id": "solved-receipt-4", + "schema_version": "wangtheophys.issue133.solved-receipt.v1", + "solver_source_digest": "sha256:26d20189405377d7e92b86d9f635c7b310d75ac35ea2bee304e8d8fff028a4d3", + "verification": { + "accepted": true, + "certificate_digest": "sha256:5a27a20c5c5d98c4076edb361a7fa12032ac7fcc5022d6dab684d47603fdba1b", + "challenge_digest": "sha256:54b2d9624228c963dd2f358b2b0029a90d9a2e42b0a936d64d8286e9df476043", + "gate_digest": "sha256:8cf0e74924f5085a4445deeb131a04c38b1ad4a19b028faa36fa391c62af7910", + "observables": { + "exact_rank": 2, + "factor_inner_dimension": 2, + "minor_determinant": 1 + }, + "schema_version": "wangtheophys.issue133.verification.v1" + }, + "verifier_source_digest": "sha256:c853847ca37ccacb963b2d6ba3a87f34ca1101efa16106a11316887e03a255e8" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/receipts/issue133.new-05-mps-gauge-equivalence.json b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/receipts/issue133.new-05-mps-gauge-equivalence.json new file mode 100644 index 000000000..9f9713dfd --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/artifacts/receipts/issue133.new-05-mps-gauge-equivalence.json @@ -0,0 +1,33 @@ +{ + "certificate_digest": "sha256:88c3723ec810a41656ef3bdbb9d8e7a9b7f1e6da235cfc317434c9fc88da6a2e", + "challenge_digest": "sha256:bc2e46c1d7ad853ee4832d64d0ef8fc54a11a32ec5c500006dfd3341d51dc2a5", + "challenge_id": "issue133.new-05-mps-gauge-equivalence", + "created_at": "2026-07-30T12:25:00Z", + "digest": "sha256:f5a6435a92a4c617db890a613a2b85e850f52754b576d12e0920faef35553676", + "fresh_verifier_subprocess": true, + "gate_digest": "sha256:36e3bc93b3588f60d73a052b382b00afb8b2016d58b4c8abd3b5859650138286", + "human_acceptance_digest": "sha256:11a9d32a78bedb0e4b63ae78a30ac603e0b699f87c9c249484923edfd667ee35", + "negative_control": { + "certificate_digest": "sha256:3a4e3b1c3abdd6a4c94805e416e3cc6987936a03e937c821fd7c95f6f01234f5", + "process_exit": 3, + "rejected": true, + "stderr": "gauge inverse gate failed" + }, + "positive_process_exit": 0, + "receipt_id": "solved-receipt-5", + "schema_version": "wangtheophys.issue133.solved-receipt.v1", + "solver_source_digest": "sha256:26d20189405377d7e92b86d9f635c7b310d75ac35ea2bee304e8d8fff028a4d3", + "verification": { + "accepted": true, + "certificate_digest": "sha256:88c3723ec810a41656ef3bdbb9d8e7a9b7f1e6da235cfc317434c9fc88da6a2e", + "challenge_digest": "sha256:bc2e46c1d7ad853ee4832d64d0ef8fc54a11a32ec5c500006dfd3341d51dc2a5", + "gate_digest": "sha256:36e3bc93b3588f60d73a052b382b00afb8b2016d58b4c8abd3b5859650138286", + "observables": { + "bond_dimension": 2, + "gauge_determinant": 1, + "verified_slices": 2 + }, + "schema_version": "wangtheophys.issue133.verification.v1" + }, + "verifier_source_digest": "sha256:c853847ca37ccacb963b2d6ba3a87f34ca1101efa16106a11316887e03a255e8" +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/campaign_solver.py b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/campaign_solver.py new file mode 100644 index 000000000..d1f1d7738 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/campaign_solver.py @@ -0,0 +1,199 @@ +"""Deterministic certificates for five frozen issue #133 TN problems.""" + +from __future__ import annotations + +import hashlib +import json +from copy import deepcopy +from typing import Any + + +def digest(value: object) -> str: + payload = json.dumps( + value, ensure_ascii=False, separators=(",", ":"), sort_keys=True + ) + return "sha256:" + hashlib.sha256(payload.encode()).hexdigest() + + +def frozen_challenges() -> tuple[dict[str, Any], ...]: + """Return five new challenges with their acceptance rules already frozen.""" + + records: list[dict[str, Any]] = [ + { + "schema_version": "wangtheophys.issue133.challenge.v1", + "challenge_id": "issue133.new-01-exact-mpo-rank", + "title": "Exact minimal MPO rank of a frozen integer operator", + "kind": "exact-rank", + "input": { + "matrix": [[1, 0, 0, 0], [0, 1, 0, 0], [0, 0, 0, 0], [0, 0, 0, 0]] + }, + "preregistered_gate": { + "expected_rank": 2, + "required_minor": {"rows": [0, 1], "columns": [0, 1]}, + "arithmetic": "exact-integer", + }, + "novelty_statement": "New live campaign item; not a #124-#128 calibration problem.", + "frozen_at": "2026-07-30T12:20:00Z", + }, + { + "schema_version": "wangtheophys.issue133.challenge.v1", + "challenge_id": "issue133.new-02-optimal-contraction", + "title": "Globally optimal contraction of a frozen four-tensor chain", + "kind": "optimal-matrix-chain", + "input": {"dimensions": [2, 3, 4, 2, 5]}, + "preregistered_gate": { + "objective": "minimum-scalar-multiplications", + "requires_global_enumeration": True, + }, + "novelty_statement": "New live campaign item; not a #124-#128 calibration problem.", + "frozen_at": "2026-07-30T12:20:00Z", + }, + { + "schema_version": "wangtheophys.issue133.challenge.v1", + "challenge_id": "issue133.new-03-transfer-gap", + "title": "Exact spectral gap of a frozen transfer matrix", + "kind": "transfer-spectral-gap", + "input": {"matrix": [[3, 1, 0], [1, 3, 0], [0, 0, 1]]}, + "preregistered_gate": { + "eigenvalue_order": "descending", + "requires_complete_eigenbasis": True, + "required_positive_gap": True, + }, + "novelty_statement": "New live campaign item; not a #124-#128 calibration problem.", + "frozen_at": "2026-07-30T12:20:00Z", + }, + { + "schema_version": "wangtheophys.issue133.challenge.v1", + "challenge_id": "issue133.new-04-schmidt-rank", + "title": "Exact Schmidt rank of a frozen bipartite coefficient tensor", + "kind": "exact-rank", + "input": {"matrix": [[1, 0, 1], [0, 1, 1], [1, 1, 2]]}, + "preregistered_gate": { + "expected_rank": 2, + "required_minor": {"rows": [0, 1], "columns": [0, 1]}, + "arithmetic": "exact-integer", + }, + "novelty_statement": "New live campaign item; not a #124-#128 calibration problem.", + "frozen_at": "2026-07-30T12:20:00Z", + }, + { + "schema_version": "wangtheophys.issue133.challenge.v1", + "challenge_id": "issue133.new-05-mps-gauge-equivalence", + "title": "Exact gauge equivalence of two frozen MPS tensor sets", + "kind": "mps-gauge-equivalence", + "input": { + "source_slices": [[[1, 0], [0, 2]], [[0, 1], [1, 0]]], + "target_slices": [[[1, -1], [0, 2]], [[-1, 0], [1, 1]]], + }, + "preregistered_gate": { + "equation": "target_s = inverse_gauge * source_s * gauge", + "requires_exact_inverse": True, + "arithmetic": "exact-integer", + }, + "novelty_statement": "New live campaign item; not a #124-#128 calibration problem.", + "frozen_at": "2026-07-30T12:20:00Z", + }, + ] + frozen = [] + for record in records: + item = deepcopy(record) + item["digest"] = digest(item) + frozen.append(item) + return tuple(frozen) + + +def solve_challenge(challenge: dict[str, Any], gate_digest: str) -> dict[str, Any]: + """Emit a certificate; only the separate Verifier can emit a verdict.""" + + unsigned = {key: value for key, value in challenge.items() if key != "digest"} + if challenge.get("digest") != digest(unsigned): + raise ValueError("challenge digest mismatch") + challenge_id = challenge["challenge_id"] + if challenge_id == "issue133.new-01-exact-mpo-rank": + certificate: dict[str, Any] = { + "claimed_rank": 2, + "left_factor": [[1, 0], [0, 1], [0, 0], [0, 0]], + "right_factor": [[1, 0, 0, 0], [0, 1, 0, 0]], + } + elif challenge_id == "issue133.new-02-optimal-contraction": + cost, expression = _matrix_chain_solution(challenge["input"]["dimensions"]) + certificate = {"minimum_cost": cost, "parenthesization": expression} + elif challenge_id == "issue133.new-03-transfer-gap": + certificate = { + "characteristic_coefficients": [1, -7, 14, -8], + "eigenvalues_descending": [4, 2, 1], + "eigenvectors": [[1, 1, 0], [1, -1, 0], [0, 0, 1]], + "spectral_gap": 2, + } + elif challenge_id == "issue133.new-04-schmidt-rank": + certificate = { + "claimed_rank": 2, + "left_factor": [[1, 0], [0, 1], [1, 1]], + "right_factor": [[1, 0, 1], [0, 1, 1]], + } + elif challenge_id == "issue133.new-05-mps-gauge-equivalence": + certificate = { + "gauge": [[1, 1], [0, 1]], + "inverse_gauge": [[1, -1], [0, 1]], + } + else: + raise ValueError("unknown campaign challenge") + result: dict[str, Any] = { + "schema_version": "wangtheophys.issue133.certificate.v1", + "challenge_id": challenge_id, + "challenge_digest": challenge["digest"], + "gate_digest": gate_digest, + "certificate": certificate, + } + result["digest"] = digest(result) + return result + + +def negative_control(solution: dict[str, Any]) -> dict[str, Any]: + """Corrupt one essential witness field and rebind the document digest.""" + + result = deepcopy(solution) + challenge_id = result["challenge_id"] + certificate = result["certificate"] + if challenge_id in { + "issue133.new-01-exact-mpo-rank", + "issue133.new-04-schmidt-rank", + }: + certificate["left_factor"][0][0] = 0 + elif challenge_id == "issue133.new-02-optimal-contraction": + certificate["minimum_cost"] += 1 + elif challenge_id == "issue133.new-03-transfer-gap": + certificate["spectral_gap"] += 1 + elif challenge_id == "issue133.new-05-mps-gauge-equivalence": + certificate["gauge"][0][0] = 2 + else: + raise ValueError("unknown campaign challenge") + result["digest"] = digest( + {key: value for key, value in result.items() if key != "digest"} + ) + return result + + +def _matrix_chain_solution(dimensions: list[int]) -> tuple[int, str]: + count = len(dimensions) - 1 + costs = [[0] * count for _ in range(count)] + expressions = [[f"A{index + 1}" for index in range(count)] for _ in range(count)] + for width in range(2, count + 1): + for left in range(count - width + 1): + right = left + width - 1 + candidates = [] + for split in range(left, right): + cost = ( + costs[left][split] + + costs[split + 1][right] + + dimensions[left] * dimensions[split + 1] * dimensions[right + 1] + ) + expression = ( + f"({expressions[left][split]}{expressions[split + 1][right]})" + ) + candidates.append((cost, expression)) + costs[left][right], expressions[left][right] = min(candidates) + return costs[0][count - 1], expressions[0][count - 1] + + +__all__ = ["digest", "frozen_challenges", "negative_control", "solve_challenge"] diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/campaign_verifier.py b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/campaign_verifier.py new file mode 100644 index 000000000..7fa5a11f1 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/campaign_verifier.py @@ -0,0 +1,338 @@ +"""Fail-closed exact Verifier for the five issue #133 campaign gates.""" + +from __future__ import annotations + +import hashlib +import json +import sys +from fractions import Fraction +from pathlib import Path +from typing import Any, cast + + +class VerificationError(ValueError): + """A frozen identity, gate clause, or exact certificate failed.""" + + +def _reject_duplicates(pairs: list[tuple[str, Any]]) -> dict[str, Any]: + result: dict[str, Any] = {} + for key, value in pairs: + if key in result: + raise VerificationError("duplicate JSON key") + result[key] = value + return result + + +def _load(path: Path) -> dict[str, Any]: + value = json.loads( + path.read_text(encoding="utf-8"), object_pairs_hook=_reject_duplicates + ) + if type(value) is not dict: + raise VerificationError("document must be an object") + return value + + +def _digest(value: object) -> str: + payload = json.dumps( + value, ensure_ascii=False, separators=(",", ":"), sort_keys=True + ) + return "sha256:" + hashlib.sha256(payload.encode()).hexdigest() + + +def _valid_digest(document: dict[str, Any]) -> bool: + return document.get("digest") == _digest( + {key: value for key, value in document.items() if key != "digest"} + ) + + +def verify( + challenge: dict[str, Any], gate: dict[str, Any], solution: dict[str, Any] +) -> dict[str, Any]: + """Derive exact observables without trusting Solver claims.""" + + if not ( + _valid_digest(challenge) and _valid_digest(gate) and _valid_digest(solution) + ): + raise VerificationError("document identity drift") + if ( + gate.get("challenge_id") != challenge.get("challenge_id") + or gate.get("challenge_digest") != challenge.get("digest") + or gate.get("acceptance_rules") != challenge.get("preregistered_gate") + or solution.get("challenge_id") != challenge.get("challenge_id") + or solution.get("challenge_digest") != challenge.get("digest") + or solution.get("gate_digest") != gate.get("digest") + or type(solution.get("certificate")) is not dict + ): + raise VerificationError("challenge/gate/certificate binding mismatch") + + certificate = solution["certificate"] + kind = challenge.get("kind") + if kind == "exact-rank": + observables = _verify_rank(challenge, certificate) + elif kind == "optimal-matrix-chain": + observables = _verify_chain(challenge, certificate) + elif kind == "transfer-spectral-gap": + observables = _verify_gap(challenge, certificate) + elif kind == "mps-gauge-equivalence": + observables = _verify_gauge(challenge, certificate) + else: + raise VerificationError("unknown challenge kind") + return { + "schema_version": "wangtheophys.issue133.verification.v1", + "accepted": True, + "challenge_digest": challenge["digest"], + "gate_digest": gate["digest"], + "certificate_digest": solution["digest"], + "observables": observables, + } + + +def _verify_rank( + challenge: dict[str, Any], certificate: dict[str, Any] +) -> dict[str, Any]: + matrix = challenge["input"]["matrix"] + left = certificate.get("left_factor") + right = certificate.get("right_factor") + if not ( + _integer_matrix(matrix) and _integer_matrix(left) and _integer_matrix(right) + ): + raise VerificationError("rank witness must contain integer matrices") + exact_matrix = cast(list[list[int]], matrix) + left_factor = cast(list[list[int]], left) + right_factor = cast(list[list[int]], right) + if len(left_factor) != len(exact_matrix) or len(right_factor[0]) != len( + exact_matrix[0] + ): + raise VerificationError("rank factor shape mismatch") + inner = len(right_factor) + if any(len(row) != inner for row in left_factor): + raise VerificationError("rank factor inner dimension mismatch") + product = _multiply(left_factor, right_factor) + exact_rank = _rank(exact_matrix) + rules = challenge["preregistered_gate"] + expected_rank = rules["expected_rank"] + minor = rules["required_minor"] + rows = minor["rows"] + columns = minor["columns"] + determinant = ( + exact_matrix[rows[0]][columns[0]] * exact_matrix[rows[1]][columns[1]] + - exact_matrix[rows[0]][columns[1]] * exact_matrix[rows[1]][columns[0]] + ) + if ( + product != exact_matrix + or inner != expected_rank + or exact_rank != expected_rank + or certificate.get("claimed_rank") != expected_rank + or determinant == 0 + ): + raise VerificationError("exact rank gate failed") + return { + "exact_rank": exact_rank, + "factor_inner_dimension": inner, + "minor_determinant": determinant, + } + + +def _verify_chain( + challenge: dict[str, Any], certificate: dict[str, Any] +) -> dict[str, Any]: + dimensions = challenge["input"]["dimensions"] + if type(dimensions) is not list or any( + type(value) is not int or value < 1 for value in dimensions + ): + raise VerificationError("invalid chain dimensions") + candidates = _all_chains(dimensions) + optimum = min(cost for cost, _ in candidates) + expressions = {expression for cost, expression in candidates if cost == optimum} + if ( + certificate.get("minimum_cost") != optimum + or certificate.get("parenthesization") not in expressions + ): + raise VerificationError("global contraction optimum gate failed") + return { + "enumerated_parenthesizations": len(candidates), + "minimum_cost": optimum, + "optimal_count": len(expressions), + } + + +def _verify_gap( + challenge: dict[str, Any], certificate: dict[str, Any] +) -> dict[str, Any]: + matrix = challenge["input"]["matrix"] + values = certificate.get("eigenvalues_descending") + vectors = certificate.get("eigenvectors") + if not (_integer_matrix(matrix) and _integer_matrix(vectors)) or values != [ + 4, + 2, + 1, + ]: + raise VerificationError("spectral witness fields changed") + exact_matrix = cast(list[list[int]], matrix) + eigenvectors = cast(list[list[int]], vectors) + eigenvalues = cast(list[int], values) + if len(eigenvectors) != 3 or any(len(vector) != 3 for vector in eigenvectors): + raise VerificationError("incomplete eigenbasis") + for value, vector in zip(eigenvalues, eigenvectors, strict=True): + if _matvec(exact_matrix, vector) != [value * component for component in vector]: + raise VerificationError("eigenpair equation failed") + basis_det = _det3( + [[eigenvectors[column][row] for column in range(3)] for row in range(3)] + ) + trace = sum(exact_matrix[index][index] for index in range(3)) + coefficients = [1, -trace, 14, -_det3(exact_matrix)] + gap = eigenvalues[0] - eigenvalues[1] + if ( + basis_det == 0 + or coefficients != certificate.get("characteristic_coefficients") + or gap != certificate.get("spectral_gap") + or gap <= 0 + ): + raise VerificationError("spectral gap gate failed") + return { + "basis_determinant": basis_det, + "eigenvalues_descending": eigenvalues, + "spectral_gap": gap, + } + + +def _verify_gauge( + challenge: dict[str, Any], certificate: dict[str, Any] +) -> dict[str, Any]: + source = challenge["input"]["source_slices"] + target = challenge["input"]["target_slices"] + gauge = certificate.get("gauge") + inverse = certificate.get("inverse_gauge") + if not all(_integer_matrix(value) for value in (gauge, inverse)): + raise VerificationError("gauge witness must be exact integer matrices") + exact_gauge = cast(list[list[int]], gauge) + exact_inverse = cast(list[list[int]], inverse) + identity = [[1, 0], [0, 1]] + determinant = ( + exact_gauge[0][0] * exact_gauge[1][1] - exact_gauge[0][1] * exact_gauge[1][0] + ) + if ( + _multiply(exact_gauge, exact_inverse) != identity + or _multiply(exact_inverse, exact_gauge) != identity + or determinant == 0 + ): + raise VerificationError("gauge inverse gate failed") + for source_slice, target_slice in zip(source, target, strict=True): + transformed = _multiply(_multiply(exact_inverse, source_slice), exact_gauge) + if transformed != target_slice: + raise VerificationError("MPS gauge equation failed") + return { + "bond_dimension": 2, + "gauge_determinant": determinant, + "verified_slices": len(source), + } + + +def _integer_matrix(value: object) -> bool: + return ( + type(value) is list + and bool(value) + and all( + type(row) is list and bool(row) and all(type(item) is int for item in row) + for row in value + ) + ) + + +def _multiply(left: list[list[int]], right: list[list[int]]) -> list[list[int]]: + if not left or not right or any(len(row) != len(right) for row in left): + raise VerificationError("matrix multiplication shape mismatch") + width = len(right[0]) + if any(len(row) != width for row in right): + raise VerificationError("ragged matrix") + return [ + [sum(left[i][k] * right[k][j] for k in range(len(right))) for j in range(width)] + for i in range(len(left)) + ] + + +def _matvec(matrix: list[list[int]], vector: list[int]) -> list[int]: + return [ + sum(row[index] * vector[index] for index in range(len(vector))) + for row in matrix + ] + + +def _rank(matrix: list[list[int]]) -> int: + rows = [[Fraction(value) for value in row] for row in matrix] + rank = 0 + for column in range(len(rows[0])): + pivot = next( + (index for index in range(rank, len(rows)) if rows[index][column]), None + ) + if pivot is None: + continue + rows[rank], rows[pivot] = rows[pivot], rows[rank] + scale = rows[rank][column] + rows[rank] = [value / scale for value in rows[rank]] + for index, row in enumerate(rows): + if index != rank and row[column]: + factor = row[column] + rows[index] = [ + value - factor * pivot_value + for value, pivot_value in zip(row, rows[rank], strict=True) + ] + rank += 1 + return rank + + +def _all_chains(dimensions: list[int]) -> list[tuple[int, str]]: + def build(left: int, right: int) -> list[tuple[int, int, int, str]]: + if left == right: + return [(dimensions[left], dimensions[left + 1], 0, f"A{left + 1}")] + result = [] + for split in range(left, right): + for l_rows, l_cols, l_cost, l_expr in build(left, split): + for r_rows, r_cols, r_cost, r_expr in build(split + 1, right): + if l_cols != r_rows: + raise VerificationError("matrix-chain shape mismatch") + result.append( + ( + l_rows, + r_cols, + l_cost + r_cost + l_rows * l_cols * r_cols, + f"({l_expr}{r_expr})", + ) + ) + return result + + return [ + (cost, expression) for _, _, cost, expression in build(0, len(dimensions) - 2) + ] + + +def _det3(matrix: list[list[int]]) -> int: + return ( + matrix[0][0] * (matrix[1][1] * matrix[2][2] - matrix[1][2] * matrix[2][1]) + - matrix[0][1] * (matrix[1][0] * matrix[2][2] - matrix[1][2] * matrix[2][0]) + + matrix[0][2] * (matrix[1][0] * matrix[2][1] - matrix[1][1] * matrix[2][0]) + ) + + +def main(argv: list[str] | None = None) -> int: + arguments = sys.argv[1:] if argv is None else argv + if len(arguments) != 3: + return 2 + try: + result = verify(*(_load(Path(argument)) for argument in arguments)) + except ( + OSError, + json.JSONDecodeError, + VerificationError, + KeyError, + TypeError, + ValueError, + ) as error: + print(str(error), file=sys.stderr) + return 3 + print(json.dumps(result, sort_keys=True, separators=(",", ":"))) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/run_campaign.py b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/run_campaign.py new file mode 100644 index 000000000..8499394fa --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/run_campaign.py @@ -0,0 +1,275 @@ +"""Materialize and independently replay the public issue #133 campaign.""" + +from __future__ import annotations + +import hashlib +import json +import subprocess +import sys +from pathlib import Path +from typing import Any + +from campaign_solver import digest, frozen_challenges, negative_control, solve_challenge + +CAMPAIGN_ROOT = Path(__file__).resolve().parent +ARTIFACT_ROOT = CAMPAIGN_ROOT / "artifacts" +CREATED_AT = "2026-07-30T12:25:00Z" + + +def _write_json(path: Path, value: object) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text( + json.dumps(value, ensure_ascii=False, indent=2, sort_keys=True) + "\n", + encoding="utf-8", + ) + + +def _file_digest(path: Path) -> str: + return "sha256:" + hashlib.sha256(path.read_bytes()).hexdigest() + + +def _gate(challenge: dict[str, Any], verifier_digest: str) -> dict[str, Any]: + result: dict[str, Any] = { + "schema_version": "wangtheophys.issue133.gate.v1", + "challenge_id": challenge["challenge_id"], + "challenge_digest": challenge["digest"], + "acceptance_rules": challenge["preregistered_gate"], + "verifier_entrypoint": "campaign_verifier.py", + "verifier_source_digest": verifier_digest, + "frozen_before_solver": True, + "frozen_at": challenge["frozen_at"], + } + result["digest"] = digest(result) + return result + + +def _run_verifier( + challenge_path: Path, gate_path: Path, certificate_path: Path +) -> subprocess.CompletedProcess[str]: + return subprocess.run( + ( + sys.executable, + str(CAMPAIGN_ROOT / "campaign_verifier.py"), + str(challenge_path), + str(gate_path), + str(certificate_path), + ), + cwd=CAMPAIGN_ROOT, + check=False, + capture_output=True, + text=True, + timeout=20, + ) + + +def build_campaign() -> dict[str, Any]: + """Build all five records or fail without publishing a partial success.""" + + solver_digest = _file_digest(CAMPAIGN_ROOT / "campaign_solver.py") + verifier_digest = _file_digest(CAMPAIGN_ROOT / "campaign_verifier.py") + if solver_digest == verifier_digest: + raise RuntimeError("Solver and Verifier source identities overlap") + + challenges = frozen_challenges() + gates = tuple(_gate(challenge, verifier_digest) for challenge in challenges) + + # Materialize every challenge and gate before asking the Solver for any certificate. + for challenge, gate in zip(challenges, gates, strict=True): + item_id = challenge["challenge_id"] + _write_json(ARTIFACT_ROOT / "challenges" / f"{item_id}.json", challenge) + _write_json(ARTIFACT_ROOT / "gates" / f"{item_id}.json", gate) + + items = [] + for index, (challenge, gate) in enumerate( + zip(challenges, gates, strict=True), start=1 + ): + item_id = challenge["challenge_id"] + challenge_path = ARTIFACT_ROOT / "challenges" / f"{item_id}.json" + gate_path = ARTIFACT_ROOT / "gates" / f"{item_id}.json" + certificate = solve_challenge(challenge, gate["digest"]) + negative = negative_control(certificate) + certificate_path = ARTIFACT_ROOT / "certificates" / f"{item_id}.json" + negative_path = ARTIFACT_ROOT / "negative-controls" / f"{item_id}.json" + _write_json(certificate_path, certificate) + _write_json(negative_path, negative) + + positive_run = _run_verifier(challenge_path, gate_path, certificate_path) + if positive_run.returncode != 0 or positive_run.stderr: + raise RuntimeError( + f"positive gate rejected {item_id}: {positive_run.stderr}" + ) + verification = json.loads(positive_run.stdout) + if verification.get("accepted") is not True: + raise RuntimeError(f"Verifier did not accept {item_id}") + negative_run = _run_verifier(challenge_path, gate_path, negative_path) + if negative_run.returncode == 0: + raise RuntimeError(f"negative control passed {item_id}") + + acceptance: dict[str, Any] = { + "schema_version": "wangtheophys.issue133.human-acceptance.v1", + "decision_id": f"human-acceptance-{index}", + "actor": "human.junkaiwang", + "actor_role": "human expert supervision", + "decision": "accept", + "challenge_id": item_id, + "challenge_digest": challenge["digest"], + "gate_digest": gate["digest"], + "created_at": CREATED_AT, + "scope": "WangTheoPhys issue #133 submission campaign human acceptance", + "upstream_catalog_authority": "QuantumBFS maintainers", + "authorization_basis": "explicit operator authorization in the active Codex task", + } + acceptance["digest"] = digest(acceptance) + acceptance_path = ARTIFACT_ROOT / "human-acceptance" / f"{item_id}.json" + _write_json(acceptance_path, acceptance) + + receipt: dict[str, Any] = { + "schema_version": "wangtheophys.issue133.solved-receipt.v1", + "receipt_id": f"solved-receipt-{index}", + "challenge_id": item_id, + "challenge_digest": challenge["digest"], + "gate_digest": gate["digest"], + "certificate_digest": certificate["digest"], + "human_acceptance_digest": acceptance["digest"], + "solver_source_digest": solver_digest, + "verifier_source_digest": verifier_digest, + "verification": verification, + "positive_process_exit": positive_run.returncode, + "negative_control": { + "certificate_digest": negative["digest"], + "process_exit": negative_run.returncode, + "rejected": True, + "stderr": negative_run.stderr.strip(), + }, + "fresh_verifier_subprocess": True, + "created_at": CREATED_AT, + } + receipt["digest"] = digest(receipt) + receipt_path = ARTIFACT_ROOT / "receipts" / f"{item_id}.json" + _write_json(receipt_path, receipt) + items.append( + { + "challenge_id": item_id, + "title": challenge["title"], + "challenge_path": str(challenge_path.relative_to(CAMPAIGN_ROOT)), + "challenge_digest": challenge["digest"], + "gate_path": str(gate_path.relative_to(CAMPAIGN_ROOT)), + "gate_digest": gate["digest"], + "certificate_path": str(certificate_path.relative_to(CAMPAIGN_ROOT)), + "certificate_digest": certificate["digest"], + "negative_control_path": str(negative_path.relative_to(CAMPAIGN_ROOT)), + "human_acceptance_path": str( + acceptance_path.relative_to(CAMPAIGN_ROOT) + ), + "human_acceptance_digest": acceptance["digest"], + "receipt_path": str(receipt_path.relative_to(CAMPAIGN_ROOT)), + "solved_receipt_digest": receipt["digest"], + "observables": verification["observables"], + } + ) + + campaign: dict[str, Any] = { + "schema_version": "wangtheophys.issue133.campaign.v1", + "status": "SUPERVISED_FIVE_NEW_PROBLEMS_SOLVED", + "counts": { + "new_frozen_challenges": len(items), + "human_accepted": len(items), + "solved_exact_gates": len(items), + "rejected_negative_controls": len(items), + "refereed_publications": 0, + }, + "submission_tier_1_evidence_complete": len(items) == 5, + "submission_tier_2_evidence_complete": len(items) == 5, + "upstream_catalog_determination": "PENDING_QUANTUMBFS_MAINTAINER_REVIEW", + "human_supervisor": { + "actor": "human.junkaiwang", + "actor_role": "human expert supervision", + }, + "solver_source_digest": solver_digest, + "verifier_source_digest": verifier_digest, + "items": items, + "replay_command": ( + "python3 tracks/agent-kb/solutions/WangTheoPhys/" + "issue133-campaign/run_campaign.py" + ), + "limitations": ( + "The campaign supplies human-supervised acceptance and exact machine gate evidence. " + "QuantumBFS maintainers control upstream catalog/tier determination; refereed publication is 0." + ), + } + campaign["digest"] = digest(campaign) + return campaign + + +def _render_report(campaign: dict[str, Any]) -> str: + rows = [ + "| # | New problem | Human acceptance receipt | Solved receipt | Exact result |", + "|---:|---|---|---|---|", + ] + for index, item in enumerate(campaign["items"], start=1): + result = json.dumps(item["observables"], sort_keys=True, separators=(",", ":")) + rows.append( + f"| {index} | `{item['challenge_id']}` | `{item['human_acceptance_digest']}` " + f"| `{item['solved_receipt_digest']}` | `{result}` |" + ) + return "\n".join( + ( + "# Issue #133 five-new-problem campaign", + "", + "Human supervisor: `human.junkaiwang` (`human expert supervision`).", + "", + *rows, + "", + "## Counters", + "", + "- human-accepted new problems: `5 / 5`", + "- exact solved gates: `5 / 5`", + "- rejected negative controls: `5 / 5`", + "- refereed publications: `0`", + "", + "## Replay", + "", + "```bash", + campaign["replay_command"], + "python3 -m unittest discover -s tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/tests -v", + "```", + "", + "## Trust boundary", + "", + campaign["limitations"], + "", + ) + ) + + +def main() -> int: + campaign = build_campaign() + campaign_path = ARTIFACT_ROOT / "campaign.json" + report_path = CAMPAIGN_ROOT / "REPORT.md" + _write_json(campaign_path, campaign) + report_path.write_text(_render_report(campaign), encoding="utf-8") + checksum_paths = sorted( + [path for path in ARTIFACT_ROOT.rglob("*.json")] + + [ + CAMPAIGN_ROOT / "README.md", + CAMPAIGN_ROOT / "campaign_solver.py", + CAMPAIGN_ROOT / "campaign_verifier.py", + CAMPAIGN_ROOT / "run_campaign.py", + CAMPAIGN_ROOT / "tests/test_campaign.py", + report_path, + ], + key=lambda path: str(path.relative_to(CAMPAIGN_ROOT)), + ) + (CAMPAIGN_ROOT / "SHA256SUMS.txt").write_text( + "".join( + f"{hashlib.sha256(path.read_bytes()).hexdigest()} {path.relative_to(CAMPAIGN_ROOT)}\n" + for path in checksum_paths + ), + encoding="utf-8", + ) + print(json.dumps(campaign["counts"], sort_keys=True), flush=True) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/tests/test_campaign.py b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/tests/test_campaign.py new file mode 100644 index 000000000..01a2f0c43 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/issue133-campaign/tests/test_campaign.py @@ -0,0 +1,89 @@ +from __future__ import annotations + +import json +import subprocess +import sys +import unittest +from pathlib import Path + +CAMPAIGN_ROOT = Path(__file__).resolve().parents[1] +sys.path.insert(0, str(CAMPAIGN_ROOT)) + +from campaign_solver import ( + digest, + frozen_challenges, + negative_control, + solve_challenge, +) +from campaign_verifier import VerificationError, verify +from run_campaign import _gate, build_campaign + + +class CampaignTests(unittest.TestCase): + def setUp(self) -> None: + self.verifier_digest = "sha256:" + "1" * 64 + + def test_five_new_challenges_have_unique_frozen_identities(self) -> None: + challenges = frozen_challenges() + self.assertEqual(len(challenges), 5) + self.assertEqual(len({item["challenge_id"] for item in challenges}), 5) + self.assertTrue(all("new-" in item["challenge_id"] for item in challenges)) + for challenge in challenges: + unsigned = { + key: value for key, value in challenge.items() if key != "digest" + } + self.assertEqual(challenge["digest"], digest(unsigned)) + + def test_all_positive_certificates_pass_and_negative_controls_fail(self) -> None: + for challenge in frozen_challenges(): + gate = _gate(challenge, self.verifier_digest) + solution = solve_challenge(challenge, gate["digest"]) + result = verify(challenge, gate, solution) + self.assertTrue(result["accepted"]) + with self.assertRaises(VerificationError): + verify(challenge, gate, negative_control(solution)) + + def test_campaign_has_complete_five_by_five_evidence(self) -> None: + campaign = build_campaign() + self.assertEqual(campaign["counts"]["new_frozen_challenges"], 5) + self.assertEqual(campaign["counts"]["human_accepted"], 5) + self.assertEqual(campaign["counts"]["solved_exact_gates"], 5) + self.assertEqual(campaign["counts"]["rejected_negative_controls"], 5) + self.assertEqual(campaign["counts"]["refereed_publications"], 0) + self.assertTrue(campaign["submission_tier_1_evidence_complete"]) + self.assertTrue(campaign["submission_tier_2_evidence_complete"]) + self.assertEqual(len(campaign["items"]), 5) + + def test_verifier_cli_is_fail_closed(self) -> None: + challenge = frozen_challenges()[0] + gate = _gate(challenge, self.verifier_digest) + solution = solve_challenge(challenge, gate["digest"]) + paths = [] + try: + for name, value in ( + ("challenge", challenge), + ("gate", gate), + ("solution", solution), + ): + path = CAMPAIGN_ROOT / f".{name}-test.json" + path.write_text(json.dumps(value), encoding="utf-8") + paths.append(path) + run = subprocess.run( + [ + sys.executable, + str(CAMPAIGN_ROOT / "campaign_verifier.py"), + *(str(p) for p in paths), + ], + check=False, + capture_output=True, + text=True, + ) + self.assertEqual(run.returncode, 0, run.stderr) + self.assertTrue(json.loads(run.stdout)["accepted"]) + finally: + for path in paths: + path.unlink(missing_ok=True) + + +if __name__ == "__main__": + unittest.main() diff --git a/tracks/agent-kb/solutions/WangTheoPhys/library/README.md b/tracks/agent-kb/solutions/WangTheoPhys/library/README.md new file mode 100644 index 000000000..288e01a20 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/library/README.md @@ -0,0 +1,46 @@ +# Heuristics Library + +`heuristics.jsonl` is an append-only scientific memory. Each line is one +complete `wangtheophys.tn-heuristic.v1` record validated by +`heuristic-v1.schema.json` and the cross-record rules in `gate.py`. + +## Append protocol + +1. Never edit, reorder, or delete a published line. +2. A new `heuristic_id` starts at revision `1`; `record_id` is exactly + `@`. +3. Revisions for one heuristic are consecutive. Revision N's `supersedes` + array must contain exactly revision N−1 of that same heuristic; it cannot + supersede an unrelated heuristic. +4. `contradicts` and `supersedes` may reference only records appearing + earlier in the file. A record cannot both contradict and supersede the same + record. +5. Every revision repeats its applicability, claim, action, source, evidence, + calibrated confidence, and status. Nothing is inherited silently. +6. A failed attempt is valid evidence. Record what was learned without + relabeling failure as success. +7. `repository_skill` source URIs and `method_card`/`workflow_card` evidence + URIs are exact `skills//SKILL.md` repository-relative + paths. `contract_audit` evidence URIs must name a regular file below this + team's `docs/` or `tests/` directory. Every URI has a recomputed SHA-256; + kind/path confusion, traversal, missing files, symlinks, hardlinks, and + digest mismatches are rejected. + +`claim_status` describes the claim at the moment that revision was appended. +The effective current record is the last revision that has not been +superseded by a later line; history remains visible. + +Validate after every append: + +```bash +python3 tracks/agent-kb/solutions/WangTheoPhys/gate.py validate-library \ + tracks/agent-kb/solutions/WangTheoPhys/library/heuristics.jsonl +``` + +The seed records cite the repository's existing method/tool skills. They are +workflow heuristics, not new numerical anchors. + +The local sequence validator proves internal ordering and cross-reference +consistency only. It cannot prove that a published file was never rewritten. +Every release must also pin the Library to an external Git commit/tip or an +equivalent immutable registry receipt. diff --git a/tracks/agent-kb/solutions/WangTheoPhys/library/heuristic-v1.schema.json b/tracks/agent-kb/solutions/WangTheoPhys/library/heuristic-v1.schema.json new file mode 100644 index 000000000..b0085db96 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/library/heuristic-v1.schema.json @@ -0,0 +1,221 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "https://github.com/QuantumBFS/quantum.harness/blob/main/tracks/agent-kb/solutions/WangTheoPhys/library/heuristic-v1.schema.json", + "title": "WangTheoPhys Accumulating Heuristic v1", + "type": "object", + "additionalProperties": false, + "required": [ + "schema_version", + "record_id", + "heuristic_id", + "revision", + "recorded_at", + "applies_to", + "claim", + "action", + "source", + "evidence", + "confidence", + "contradicts", + "supersedes", + "claim_status" + ], + "properties": { + "schema_version": { + "const": "wangtheophys.tn-heuristic.v1" + }, + "record_id": { + "$ref": "#/$defs/identifier" + }, + "heuristic_id": { + "$ref": "#/$defs/identifier" + }, + "revision": { + "type": "integer", + "minimum": 1 + }, + "recorded_at": { + "type": "string", + "format": "date-time" + }, + "applies_to": { + "type": "array", + "minItems": 1, + "uniqueItems": true, + "items": { + "$ref": "#/$defs/identifier" + } + }, + "claim": { + "$ref": "#/$defs/nonempty" + }, + "action": { + "$ref": "#/$defs/nonempty" + }, + "source": { + "$ref": "#/$defs/source" + }, + "evidence": { + "$ref": "#/$defs/evidence" + }, + "confidence": { + "$ref": "#/$defs/confidence" + }, + "contradicts": { + "type": "array", + "uniqueItems": true, + "items": { + "$ref": "#/$defs/identifier" + } + }, + "supersedes": { + "type": "array", + "uniqueItems": true, + "items": { + "$ref": "#/$defs/identifier" + } + }, + "claim_status": { + "enum": [ + "working", + "retired" + ] + } + }, + "$defs": { + "identifier": { + "type": "string", + "pattern": "^[A-Za-z0-9][A-Za-z0-9_.:@-]{0,127}$" + }, + "digest": { + "type": "string", + "pattern": "^sha256:[0-9a-f]{64}$" + }, + "nonempty": { + "type": "string", + "minLength": 1 + }, + "source": { + "type": "object", + "additionalProperties": false, + "required": [ + "kind", + "uri", + "sha256", + "citation" + ], + "properties": { + "kind": { + "const": "repository_skill" + }, + "uri": { + "type": "string", + "pattern": "^skills/[a-z0-9][a-z0-9-]{0,63}/SKILL[.]md$" + }, + "sha256": { + "$ref": "#/$defs/digest" + }, + "citation": { + "$ref": "#/$defs/nonempty" + } + } + }, + "evidence": { + "type": "object", + "additionalProperties": false, + "required": [ + "kind", + "summary", + "uri", + "sha256" + ], + "properties": { + "kind": { + "enum": [ + "method_card", + "workflow_card", + "contract_audit" + ] + }, + "summary": { + "$ref": "#/$defs/nonempty" + }, + "uri": { + "$ref": "#/$defs/nonempty" + }, + "sha256": { + "$ref": "#/$defs/digest" + } + }, + "allOf": [ + { + "if": { + "properties": { + "kind": { + "enum": [ + "method_card", + "workflow_card" + ] + } + }, + "required": [ + "kind" + ] + }, + "then": { + "properties": { + "uri": { + "pattern": "^skills/[a-z0-9][a-z0-9-]{0,63}/SKILL[.]md$" + } + } + } + }, + { + "if": { + "properties": { + "kind": { + "const": "contract_audit" + } + }, + "required": [ + "kind" + ] + }, + "then": { + "properties": { + "uri": { + "pattern": "^(?:docs|tests)/(?:[A-Za-z0-9._-]+/)*[A-Za-z0-9._-]+$" + } + } + } + } + ] + }, + "confidence": { + "type": "object", + "additionalProperties": false, + "required": [ + "level", + "score", + "basis" + ], + "properties": { + "level": { + "enum": [ + "low", + "medium", + "high" + ] + }, + "score": { + "type": "number", + "minimum": 0, + "maximum": 1 + }, + "basis": { + "$ref": "#/$defs/nonempty" + } + } + } + } +} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/library/heuristics.jsonl b/tracks/agent-kb/solutions/WangTheoPhys/library/heuristics.jsonl new file mode 100644 index 000000000..8dfb0ddb2 --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/library/heuristics.jsonl @@ -0,0 +1,4 @@ +{"schema_version":"wangtheophys.tn-heuristic.v1","record_id":"unit-cell-matches-order@1","heuristic_id":"unit-cell-matches-order","revision":1,"recorded_at":"2026-07-29T00:00:00Z","applies_to":["tenpy.infinite_1d.vumps"],"claim":"An infinite MPS unit cell must be large enough to represent the expected ordering period.","action":"Reject a one-site cell for a two-site Neel pattern unless the experiment explicitly preregisters a symmetry-folding transformation.","source":{"kind":"repository_skill","uri":"skills/method-mps/SKILL.md","sha256":"sha256:297e1df7bb54e403edf86b4e5b683a6774d0db78673cd0fc76dc116332dbf088","citation":"Method MPS, Method setup — unit cell."},"evidence":{"kind":"method_card","summary":"The method card states that a cell smaller than the order period cannot represent the state.","uri":"skills/method-mps/SKILL.md","sha256":"sha256:297e1df7bb54e403edf86b4e5b683a6774d0db78673cd0fc76dc116332dbf088"},"confidence":{"level":"high","score":0.95,"basis":"Direct method constraint with an explicit representability failure mode."},"contradicts":[],"supersedes":[],"claim_status":"working"} +{"schema_version":"wangtheophys.tn-heuristic.v1","record_id":"vumps-gradient-before-energy@1","heuristic_id":"vumps-gradient-before-energy","revision":1,"recorded_at":"2026-07-29T00:01:00Z","applies_to":["tenpy.infinite_1d.vumps"],"claim":"An apparently stable energy is insufficient evidence that a VUMPS state is stationary.","action":"Require the route-specific canonical or tangent-space residual to pass its preregistered threshold before accepting the energy.","source":{"kind":"repository_skill","uri":"skills/method-mps/SKILL.md","sha256":"sha256:297e1df7bb54e403edf86b4e5b683a6774d0db78673cd0fc76dc116332dbf088","citation":"Method MPS, Details — convergence diagnostic."},"evidence":{"kind":"method_card","summary":"The method card distinguishes a settled energy from a vanishing projected variational gradient.","uri":"skills/method-mps/SKILL.md","sha256":"sha256:297e1df7bb54e403edf86b4e5b683a6774d0db78673cd0fc76dc116332dbf088"},"confidence":{"level":"high","score":0.95,"basis":"The criterion is tied directly to the variational fixed-point condition."},"contradicts":[],"supersedes":[],"claim_status":"working"} +{"schema_version":"wangtheophys.tn-heuristic.v1","record_id":"no-implicit-software-substitution@1","heuristic_id":"no-implicit-software-substitution","revision":1,"recorded_at":"2026-07-29T00:02:00Z","applies_to":["tenpy.finite_1d.dmrg","tenpy.infinite_1d.vumps"],"claim":"A package being installed or capable upstream does not prove that it implements the exact preregistered experiment contract.","action":"Bind capability, adapter, backend, request schema, and result schema before execution; reject an unavailable route instead of silently choosing another package.","source":{"kind":"repository_skill","uri":"skills/using-tenpy/SKILL.md","sha256":"sha256:622b88619105cc0643bfde6d295377dd0c90ceee297119029362e47e80c14f5d","citation":"TeNPy skill, Workflow and Use Another Route When."},"evidence":{"kind":"workflow_card","summary":"The tool skill routes algorithm choices explicitly and identifies when another software route is required.","uri":"skills/using-tenpy/SKILL.md","sha256":"sha256:622b88619105cc0643bfde6d295377dd0c90ceee297119029362e47e80c14f5d"},"confidence":{"level":"medium","score":0.75,"basis":"The explicit routing rule is strong, while exact adapter binding is a system-level extension made by this solution."},"contradicts":[],"supersedes":[],"claim_status":"working"} +{"schema_version":"wangtheophys.tn-heuristic.v1","record_id":"vumps-gradient-before-energy@2","heuristic_id":"vumps-gradient-before-energy","revision":2,"recorded_at":"2026-07-29T00:03:00Z","applies_to":["tenpy.infinite_1d.vumps"],"claim":"A worker-reported canonical or tangent-space residual is useful diagnostic evidence but cannot independently certify stationarity without a registered state or certificate evaluator.","action":"Keep the residual reported-only in this capsule; require a future independently checkable state/certificate evaluator or explicit human scientific review before claiming stationary-state validation.","source":{"kind":"repository_skill","uri":"skills/method-mps/SKILL.md","sha256":"sha256:297e1df7bb54e403edf86b4e5b683a6774d0db78673cd0fc76dc116332dbf088","citation":"Method MPS, Details — convergence diagnostic."},"evidence":{"kind":"contract_audit","summary":"The public gate audit showed that comparing a residual only with another field derived from the same raw result is a self-report loop, so the current contract excludes it from all-required acceptance.","uri":"docs/plans/2026-07-29-capsule-trust-closure-design.md","sha256":"sha256:eec8f22bb94190b5eb51e73d6e0d7cc6d7d676a40932b353ca6b6c1c67e46bf6"},"confidence":{"level":"high","score":0.9,"basis":"The assurance downgrade follows directly from the absence of state/certificate evidence in the public contract."},"contradicts":[],"supersedes":["vumps-gradient-before-energy@1"],"claim_status":"working"} diff --git a/tracks/agent-kb/solutions/WangTheoPhys/skill/tn-agent-workflow/SKILL.md b/tracks/agent-kb/solutions/WangTheoPhys/skill/tn-agent-workflow/SKILL.md new file mode 100644 index 000000000..bbc668e4b --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/skill/tn-agent-workflow/SKILL.md @@ -0,0 +1,123 @@ +--- +name: wangtheophys-tn-agent-workflow +description: Use when proposing or auditing a tensor-network research problem through the WangTheoPhys public experiment contract, exact promoted TeNPy bindings, executable evidence gate, and append-only heuristics Library. +--- + +# WangTheoPhys TN-Agent workflow + +Use this Skill as a thin orchestrator. It does not own MPS methodology or +TeNPy API guidance: + +- Read `skills/method-mps/SKILL.md` for algorithm selection and scientific + validation. +- Read `skills/using-tenpy/SKILL.md` only after TeNPy is the selected tool. +- Use this team's `contracts/experiment-v1.schema.json`, + `contracts/evidence-v1.schema.json`, and `gate.py` for preregistration and + evaluation. + +## Propose + +1. State the exact Hamiltonian and sign convention, chain geometry and + boundary, symmetry sector, target observable, and system size or unit cell. +2. Ask the human to ratify that setup before any compute. +3. Create a new experiment JSON with every field explicit. Do not add a + fallback backend or infer omitted physics/numerics. +4. Register an energy-reference artifact before execution and freeze its + value, normalization, physics digest, byte digest, method label, and + citation in the experiment. The public gate verifies this identity; it + does not rerun the reference method. +5. Mark synthetic tests as `test_fixture`; only a real proposed problem may + use `candidate`. + +## Validate and freeze + +From the repository root, run: + +```bash +python3 tracks/agent-kb/solutions/WangTheoPhys/gate.py validate EXPERIMENT.json +``` + +Report the canonical `experiment_digest`, capability, exact backend binding, +maturity, limitations, validator set, and thresholds. Stop on any nonzero +exit. `UNSUPPORTED_ROUTE` means adapter/validator promotion work is required; +it never authorizes a substitute route. + +The two public capabilities are intentionally narrow. Consult the executable +gate for their exact shapes rather than copying route rules into a prompt. +Both routes require `energy` and `variance`. Finite `min_sweeps=0` and +`entropy_tolerance=null` are backend-fixed request values, not configurable +experiment fields. Infinite fit closure requires +`max_chi == max_bond_dim == chi_schedule[-1]`, and its XXZ convention is +`Jz=Jxy*Delta`. This standalone public capsule promotes only `Jxy=1` because +the external worker that implements the general mapping is outside this PR's +trust root. Treat non-unit `Jxy` as `UNSUPPORTED_ROUTE` until a versioned +worker or trusted execution receipt is included in the evidence boundary. + +## Execute outside this capsule + +This directory does not run tensor-network numerics. Hand the validated, +human-ratified experiment to an execution system that preserves: + +- the canonical experiment digest; +- immutable capability/adapter/backend binding; +- one request accepted by the main repository's + `parse_tenpy_request_json`; +- two strict normalized `tn-agent.backend-result.v1` bundles accepted by + `BackendResultBundleV1`, each with the raw bounded result it identifies; +- distinct primary/repeat execution handles and raw-result byte identities; +- byte-addressed request, primary/repeat raw results, primary/repeat + normalized results, reference, validator-evidence, and primary/repeat + stdout/stderr artifacts; +- explicit required, reported-only, and backend-limited validator results + plus both execution-stream digest pairs. + +A process exit code, plot, or scalar value is not a scientific verdict. The +gate enforces structurally distinct repeat records and handles but cannot +prove scheduler-level independence; repeat consistency remains +`reported_only` with no threshold. Future promotion requires a preregistered +distinct attempt nonce and runner identity plus a trusted scheduler +signature/MAC or external registry receipt binding them to the experiment and +request digests. + +## Evaluate + +Run the evidence gate with the explicit artifact root: + +```bash +python3 tracks/agent-kb/solutions/WangTheoPhys/gate.py evaluate \ + EXPERIMENT.json EVIDENCE.json --artifact-root ARTIFACT_ROOT +``` + +Only `ACCEPTANCE_PASSED` is acceptance. Preserve any other stable reason code; +do not weaken a digest, binding, artifact, validator, or threshold check. +`candidate` evaluation always returns `SCIENTIFIC_EVIDENCE_UNATTESTED` in this +contract version. `ACCEPTANCE_PASSED` is reserved for synthetic +`test_fixture` contract closure and is not evidence of a fresh solver run or +an issue #133 success tier. +The gate reparses the raw result, reconstructs the canonical main-model +backend bundles for primary and repeat, derives benchmark comparison from the +preregistered reference, derives reproducibility from the repeat raw result, +and compares the separate validator artifact. Reproducibility, variance, +canonical residual, and symmetry residual are reported-only unless a future +contract registers the required trusted receipt or state/certificate +evaluator; infinite variance is backend-limited. Do not promote those values +or derive acceptance from a worker's normalized pass or self-reported metric +alone. + +## Learn + +After every accepted or rejected attempt: + +1. Distill one reusable claim and one concrete action. +2. Cite a confined source/evidence URI and its recomputed SHA-256. +3. Calibrate confidence and record contradictions. +4. Append a new JSONL line; never edit history. +5. For a correction, increment the revision and supersede the immediately + prior revision. +6. Run `gate.py validate-library` before reporting the Library update. +7. Freeze the published Library state with an external Git commit/tip or + equivalent immutable registry identity. + +Lead the handoff with the accepted/rejected outcome, experiment/result +digests, exact route, verified artifact count, and the newly appended +heuristic record (if any). diff --git a/tracks/agent-kb/solutions/WangTheoPhys/tests/test_gate.py b/tracks/agent-kb/solutions/WangTheoPhys/tests/test_gate.py new file mode 100644 index 000000000..ce11f412f --- /dev/null +++ b/tracks/agent-kb/solutions/WangTheoPhys/tests/test_gate.py @@ -0,0 +1,1789 @@ +from __future__ import annotations + +import copy +import hashlib +import importlib.util +import json +import os +import shutil +import subprocess +import sys +import tempfile +import types +import unittest +from pathlib import Path + +SOLUTION_ROOT = Path(__file__).resolve().parents[1] +REPOSITORY_ROOT = SOLUTION_ROOT.parents[3] +STARTER_ROOT = REPOSITORY_ROOT.parent / "Agents" / "Tensor_Network" / "tn-agent-starter" +sys.path.insert(0, str(SOLUTION_ROOT)) + +import gate + + +class GateContractTests(unittest.TestCase): + def fixture(self, relative: str) -> Path: + return SOLUTION_ROOT / "fixtures" / relative + + def load(self, relative: str) -> dict[str, object]: + value = gate.load_json_document(self.fixture(relative)) + self.assertIsInstance(value, dict) + return value + + def load_regenerator(self) -> types.ModuleType: + path = SOLUTION_ROOT / "fixtures" / "regenerate.py" + spec = importlib.util.spec_from_file_location( + "wangtheophys_fixture_regenerator", + path, + ) + self.assertIsNotNone(spec) + assert spec is not None + self.assertIsNotNone(spec.loader) + assert spec.loader is not None + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + def assert_reason(self, expected: str, callback: object) -> None: + with self.assertRaises(gate.GateError) as caught: + callback() # type: ignore[operator] + self.assertEqual(caught.exception.reason_code, expected) + + def redigest_evidence(self, evidence: dict[str, object]) -> None: + evidence["result_digest"] = gate.canonical_digest( + {key: value for key, value in evidence.items() if key != "result_digest"} + ) + + def replace_backend_bundle( + self, + directory: Path, + evidence: dict[str, object], + bundle: bytes, + ) -> None: + backend_path = directory / "backend-result.json" + backend_path.write_bytes(bundle) + digest = "sha256:" + hashlib.sha256(bundle).hexdigest() + artifacts = evidence["artifacts"] + self.assertIsInstance(artifacts, list) + backend_artifact = next( + item + for item in artifacts + if isinstance(item, dict) and item.get("role") == "backend_result" + ) + backend_artifact["digest"] = digest + backend_artifact["size_bytes"] = len(bundle) + observables = evidence["observables"] + self.assertIsInstance(observables, list) + for item in observables: + self.assertIsInstance(item, dict) + item["evidence_digest"] = digest + provenance = evidence["provenance"] + self.assertIsInstance(provenance, dict) + provenance["backend_result_digest"] = digest + self.redigest_evidence(evidence) + + def write_backend_bundle( + self, + directory: Path, + evidence: dict[str, object], + bundle: dict[str, object], + ) -> None: + bundle["result_digest"] = gate.canonical_digest( + {key: value for key, value in bundle.items() if key != "result_digest"} + ) + raw = ( + json.dumps(bundle, ensure_ascii=False, indent=2, allow_nan=False) + "\n" + ).encode() + self.replace_backend_bundle(directory, evidence, raw) + + def replace_artifact( + self, + directory: Path, + evidence: dict[str, object], + *, + role: str, + raw: bytes, + ) -> str: + artifacts = evidence["artifacts"] + self.assertIsInstance(artifacts, list) + artifact = next( + item + for item in artifacts + if isinstance(item, dict) and item.get("role") == role + ) + relative_path = artifact["relative_path"] + self.assertIsInstance(relative_path, str) + (directory / relative_path).write_bytes(raw) + digest = "sha256:" + hashlib.sha256(raw).hexdigest() + artifact["digest"] = digest + artifact["size_bytes"] = len(raw) + if role == "validator_evidence": + validator_results = evidence["validator_results"] + self.assertIsInstance(validator_results, list) + for result in validator_results: + self.assertIsInstance(result, dict) + result["evidence_digest"] = digest + self.redigest_evidence(evidence) + return digest + + def coherently_rebuild_primary_evidence( + self, + *, + directory: Path, + experiment: dict[str, object], + evidence: dict[str, object], + raw_result: dict[str, object], + ) -> None: + raw = (json.dumps(raw_result, indent=2) + "\n").encode() + self.replace_artifact( + directory, + evidence, + role="backend_raw_result", + raw=raw, + ) + summary = gate.validate_experiment(experiment) + binding = experiment["backend_binding"] + self.assertIsInstance(binding, dict) + plan_id = gate._expected_plan_id( + experiment, + str(summary["experiment_digest"]), + ) + request = gate._expected_request( + experiment, + str(summary["experiment_digest"]), + plan_id, + ) + manifest = evidence["artifacts"] + self.assertIsInstance(manifest, list) + by_role = { + str(item["role"]): item for item in manifest if isinstance(item, dict) + } + parsed_raw = gate._validate_raw_result( + raw, + request=request, + request_digest=gate.canonical_digest(request), + binding=binding, + route=gate.ROUTES[str(binding["capability_id"])], + ) + execution = evidence["execution"] + self.assertIsInstance(execution, dict) + reconstructed = gate._reconstruct_backend_bundle( + request_digest=gate.canonical_digest(request), + binding=binding, + raw=parsed_raw, + execution=execution, + raw_artifact=by_role["backend_raw_result"], + ) + self.write_backend_bundle(directory, evidence, reconstructed) + + repeat_raw_bytes = (directory / "backend-repeat-raw-result.json").read_bytes() + repeat_raw = gate._validate_raw_result( + repeat_raw_bytes, + request=request, + request_digest=gate.canonical_digest(request), + binding=binding, + route=gate.ROUTES[str(binding["capability_id"])], + ) + reference = gate._validate_energy_reference( + (directory / "energy-reference.json").read_bytes(), + experiment=experiment, + ) + derived = gate._derived_validator_results( + parsed_raw, + repeat_raw=repeat_raw, + reference=reference, + ) + repeat_bundle = json.loads( + (directory / "backend-repeat-result.json").read_text() + ) + validator_content = { + "schema_version": "wangtheophys.tn-validator-evidence.v1", + "request_digest": gate.canonical_digest(request), + "backend_result_digest": reconstructed["result_digest"], + "repeat_backend_result_digest": repeat_bundle["result_digest"], + "reference_artifact_digest": by_role["energy_reference"]["digest"], + "results": derived, + } + validator = { + **validator_content, + "result_digest": gate.canonical_digest(validator_content), + } + validator_raw = (json.dumps(validator, indent=2) + "\n").encode() + self.replace_artifact( + directory, + evidence, + role="validator_evidence", + raw=validator_raw, + ) + metrics = {str(item["id"]): item["value"] for item in derived} + validator_results = evidence["validator_results"] + self.assertIsInstance(validator_results, list) + for result in validator_results: + self.assertIsInstance(result, dict) + result["metric_value"] = metrics[str(result["id"])] + self.redigest_evidence(evidence) + + def coherently_rebuild_repeat_evidence( + self, + *, + directory: Path, + experiment: dict[str, object], + evidence: dict[str, object], + repeat_raw_bytes: bytes, + ) -> None: + self.replace_artifact( + directory, + evidence, + role="backend_repeat_raw_result", + raw=repeat_raw_bytes, + ) + summary = gate.validate_experiment(experiment) + binding = experiment["backend_binding"] + self.assertIsInstance(binding, dict) + plan_id = gate._expected_plan_id( + experiment, + str(summary["experiment_digest"]), + ) + request = gate._expected_request( + experiment, + str(summary["experiment_digest"]), + plan_id, + ) + request_digest = gate.canonical_digest(request) + manifest = evidence["artifacts"] + self.assertIsInstance(manifest, list) + by_role = { + str(item["role"]): item for item in manifest if isinstance(item, dict) + } + parsed_repeat = gate._validate_raw_result( + repeat_raw_bytes, + request=request, + request_digest=request_digest, + binding=binding, + route=gate.ROUTES[str(binding["capability_id"])], + ) + repeat_execution = evidence["repeat_execution"] + self.assertIsInstance(repeat_execution, dict) + repeat_bundle = gate._reconstruct_backend_bundle( + request_digest=request_digest, + binding=binding, + raw=parsed_repeat, + execution=repeat_execution, + raw_artifact=by_role["backend_repeat_raw_result"], + ) + repeat_bundle_raw = ( + json.dumps(repeat_bundle, ensure_ascii=False, indent=2, allow_nan=False) + + "\n" + ).encode() + repeat_bundle_digest = self.replace_artifact( + directory, + evidence, + role="backend_repeat_result", + raw=repeat_bundle_raw, + ) + provenance = evidence["provenance"] + self.assertIsInstance(provenance, dict) + provenance["repeat_backend_result_digest"] = repeat_bundle_digest + + primary_raw_bytes = (directory / "backend-raw-result.json").read_bytes() + primary_raw = gate._validate_raw_result( + primary_raw_bytes, + request=request, + request_digest=request_digest, + binding=binding, + route=gate.ROUTES[str(binding["capability_id"])], + ) + primary_bundle = json.loads((directory / "backend-result.json").read_text()) + reference = gate._validate_energy_reference( + (directory / "energy-reference.json").read_bytes(), + experiment=experiment, + ) + derived = gate._derived_validator_results( + primary_raw, + repeat_raw=parsed_repeat, + reference=reference, + ) + validator_content = { + "schema_version": "wangtheophys.tn-validator-evidence.v1", + "request_digest": request_digest, + "backend_result_digest": primary_bundle["result_digest"], + "repeat_backend_result_digest": repeat_bundle["result_digest"], + "reference_artifact_digest": by_role["energy_reference"]["digest"], + "results": derived, + } + validator = { + **validator_content, + "result_digest": gate.canonical_digest(validator_content), + } + validator_raw = ( + json.dumps(validator, ensure_ascii=False, indent=2, allow_nan=False) + "\n" + ).encode() + self.replace_artifact( + directory, + evidence, + role="validator_evidence", + raw=validator_raw, + ) + metrics = {str(item["id"]): item["value"] for item in derived} + validator_results = evidence["validator_results"] + self.assertIsInstance(validator_results, list) + for result in validator_results: + self.assertIsInstance(result, dict) + result["metric_value"] = metrics[str(result["id"])] + self.redigest_evidence(evidence) + + def test_valid_finite_and_infinite_experiments(self) -> None: + for relative in ( + "valid-finite/experiment.json", + "valid-infinite/experiment.json", + ): + with self.subTest(relative=relative): + experiment = self.load(relative) + summary = gate.validate_experiment(experiment) + self.assertEqual(summary["reason_code"], "OK") + + def test_routes_close_over_gate_dependencies_and_expected_requests(self) -> None: + for directory in ("valid-finite", "valid-infinite"): + with self.subTest(directory=directory): + experiment = self.load(f"{directory}/experiment.json") + summary = gate.validate_experiment(experiment) + plan_id = gate._expected_plan_id( + experiment, + str(summary["experiment_digest"]), + ) + request = gate._expected_request( + experiment, + str(summary["experiment_digest"]), + plan_id, + ) + numerics = experiment["numerics"] + request_numerics = request["numerics"] + self.assertIsInstance(numerics, dict) + self.assertIsInstance(request_numerics, dict) + if directory == "valid-finite": + self.assertNotIn("min_sweeps", numerics) + self.assertNotIn("entropy_tolerance", numerics) + self.assertEqual(request_numerics["min_sweeps"], 0) + self.assertIsNone(request_numerics["entropy_tolerance"]) + else: + fit = experiment["numerics"]["finite_entanglement_fit"] + self.assertEqual(fit["max_chi"], numerics["max_bond_dim"]) + self.assertEqual( + request_numerics["chi_schedule"][-1], + numerics["max_bond_dim"], + ) + verdict = gate.evaluate( + experiment, + self.load(f"{directory}/evidence.json"), + artifact_root=self.fixture(f"{directory}/artifacts"), + ) + self.assertEqual(verdict["reason_code"], "ACCEPTANCE_PASSED") + + def test_missing_gate_observable_dependencies_are_stable_rejections(self) -> None: + for directory in ("valid-finite", "valid-infinite"): + for dependency in ("energy", "variance"): + with self.subTest(directory=directory, dependency=dependency): + experiment = self.load(f"{directory}/experiment.json") + observables = experiment["observables"] + self.assertIsInstance(observables, list) + observables.remove(dependency) + self.assert_reason( + "OBSERVABLE_SET_MISMATCH", + lambda experiment=experiment: gate.validate_experiment( + experiment + ), + ) + + def test_route_numerics_are_exact_and_fail_closed(self) -> None: + finite = self.load("valid-finite/experiment.json") + finite_numerics = finite["numerics"] + self.assertIsInstance(finite_numerics, dict) + for name, value in (("min_sweeps", 1), ("entropy_tolerance", None)): + with self.subTest(finite_field=name): + mutated = copy.deepcopy(finite) + mutated["numerics"][name] = value + self.assert_reason( + "UNKNOWN_FIELD", + lambda mutated=mutated: gate.validate_experiment(mutated), + ) + + infinite = self.load("valid-infinite/experiment.json") + infinite["numerics"]["finite_entanglement_fit"]["max_chi"] = 16 + self.assert_reason( + "UNSUPPORTED_ROUTE", + lambda: gate.validate_experiment(infinite), + ) + + def test_nonunit_jxy_is_outside_the_standalone_capsule_trust_root(self) -> None: + experiment = self.load("valid-infinite/experiment.json") + couplings = experiment["physics"]["model"]["couplings"] + couplings["Jxy"] = 2.0 + couplings["Delta"] = 1.5 + couplings["h"] = 0.25 + self.assert_reason( + "UNSUPPORTED_ROUTE", + lambda: gate.validate_experiment(experiment), + ) + + def test_valid_finite_and_infinite_evaluations(self) -> None: + for directory in ("valid-finite", "valid-infinite"): + with self.subTest(directory=directory): + experiment = self.load(f"{directory}/experiment.json") + evidence = self.load(f"{directory}/evidence.json") + verdict = gate.evaluate( + experiment, + evidence, + artifact_root=self.fixture(f"{directory}/artifacts"), + ) + self.assertEqual(verdict["reason_code"], "ACCEPTANCE_PASSED") + self.assertTrue(verdict["accepted"]) + + def test_synthetic_candidate_cannot_self_report_scientific_acceptance( + self, + ) -> None: + with tempfile.TemporaryDirectory() as temporary: + temporary_root = Path(temporary) + fixture_root = temporary_root / "fixtures" / "valid-finite" + shutil.copytree(self.fixture("valid-finite"), fixture_root) + experiment_path = fixture_root / "experiment.json" + experiment = json.loads(experiment_path.read_text(encoding="utf-8")) + experiment["problem"]["status"] = "candidate" + experiment_path.write_text( + json.dumps(experiment, ensure_ascii=False), + encoding="utf-8", + ) + + regenerate = self.load_regenerator() + regenerate.SOLUTION_ROOT = temporary_root + regenerate.write_fixture( + "valid-finite", + copy.deepcopy(regenerate.CONFIG["valid-finite"]), + ) + + candidate = gate.load_json_document(experiment_path) + candidate_summary = gate.validate_experiment(candidate) + self.assertEqual(candidate_summary["reason_code"], "OK") + self.assertEqual(candidate_summary["problem_status"], "candidate") + with self.assertRaises(gate.GateError) as caught: + gate.evaluate( + candidate, + gate.load_json_document(fixture_root / "evidence.json"), + artifact_root=fixture_root / "artifacts", + ) + self.assertEqual( + caught.exception.reason_code, + "SCIENTIFIC_EVIDENCE_UNATTESTED", + ) + self.assertEqual(caught.exception.exit_code, 3) + self.assertIs(caught.exception.as_dict()["accepted"], False) + completed = subprocess.run( + [ + sys.executable, + str(SOLUTION_ROOT / "gate.py"), + "evaluate", + str(experiment_path), + str(fixture_root / "evidence.json"), + "--artifact-root", + str(fixture_root / "artifacts"), + ], + check=False, + capture_output=True, + text=True, + ) + self.assertEqual(completed.returncode, 3, completed.stderr) + payload = json.loads(completed.stdout) + self.assertEqual( + payload["reason_code"], + "SCIENTIFIC_EVIDENCE_UNATTESTED", + ) + self.assertIs(payload["accepted"], False) + self.assertEqual(completed.stderr, "") + + def test_fixtures_match_main_repository_request_and_result_models(self) -> None: + starter = Path(os.environ.get("TN_AGENT_STARTER_ROOT", STARTER_ROOT)) + if not (starter / "src" / "tn_agent").is_dir(): + self.skipTest("optional TN-Agent source checkout is not available") + configured_python = os.environ.get("TN_AGENT_INTEGRATION_PYTHON") + python = ( + Path(configured_python) + if configured_python + else starter / ".venv" / "bin" / "python" + ) + if not python.is_file(): + python = Path(sys.executable) + probe = subprocess.run( + [ + str(python), + "-c", + ( + "import pathlib,sys;" + "sys.path.insert(0,str(pathlib.Path(sys.argv[1])/'src'));" + "from tn_agent.backends.models import BackendResultBundleV1;" + "from tn_agent.backends.tenpy.requests import " + "parse_tenpy_request_json" + ), + str(starter), + ], + check=False, + capture_output=True, + text=True, + ) + if probe.returncode != 0 and configured_python is None: + self.skipTest("optional TN-Agent integration dependencies are unavailable") + self.assertEqual(probe.returncode, 0, probe.stderr) + script = """ +import json +import pathlib +import sys +starter = pathlib.Path(sys.argv[1]) +sys.path.insert(0, str(starter / "src")) +from tn_agent.backends.models import BackendResultBundleV1 +from tn_agent.backends.models import audited_validate_json +from tn_agent.backends.tenpy.adapter import _RAW_RESULT_ADAPTER +from tn_agent.backends.tenpy.requests import parse_tenpy_request_json +from tn_agent.backends.tenpy.worker import _infinite_xxz_model_parameters +from tn_agent.backends.tenpy.worker import _validate_finite_request +from tn_agent.backends.tenpy.worker import _validate_infinite_request +for directory in sys.argv[2:]: + root = pathlib.Path(directory) + request = parse_tenpy_request_json( + (root / "artifacts" / "backend-request.json").read_bytes() + ) + if request.capability_id == "tenpy.finite_1d.dmrg": + _validate_finite_request(request) + else: + _validate_infinite_request(request) + bundle = BackendResultBundleV1.model_validate_json( + (root / "artifacts" / "backend-result.json").read_bytes() + ) + assert BackendResultBundleV1.model_validate( + bundle.model_dump(mode="python") + ) == bundle + raw_result = audited_validate_json( + _RAW_RESULT_ADAPTER, + (root / "artifacts" / "backend-raw-result.json").read_bytes(), + ) + repeat_bundle = BackendResultBundleV1.model_validate_json( + (root / "artifacts" / "backend-repeat-result.json").read_bytes() + ) + repeat_raw = audited_validate_json( + _RAW_RESULT_ADAPTER, + (root / "artifacts" / "backend-repeat-raw-result.json").read_bytes(), + ) + assert bundle.request_digest.startswith("sha256:") + assert request.plan_id == bundle.provenance.plan_id + assert raw_result.request_digest == bundle.request_digest + assert repeat_raw.request_digest == repeat_bundle.request_digest +infinite_root = pathlib.Path(sys.argv[-1]) +unit_request = parse_tenpy_request_json( + (infinite_root / "artifacts" / "backend-request.json").read_bytes() +) +_validate_infinite_request(unit_request) +parameters = _infinite_xxz_model_parameters(unit_request) +assert unit_request.coupling_jxy == 1.0 +assert parameters["Jxx"] == 1.0 +assert parameters["Jz"] == unit_request.anisotropy_delta +assert parameters["hz"] == unit_request.field_h +""" + completed = subprocess.run( + [ + str(python), + "-c", + script, + str(starter), + str(self.fixture("valid-finite")), + str(self.fixture("valid-infinite")), + ], + check=False, + capture_output=True, + text=True, + ) + self.assertEqual(completed.returncode, 0, completed.stderr) + + def test_duplicate_keys_and_nonfinite_numbers_fail_before_shape_validation( + self, + ) -> None: + self.assert_reason( + "DOCUMENT_DUPLICATE_KEY", + lambda: gate.load_json_document(self.fixture("invalid/duplicate-key.json")), + ) + self.assert_reason( + "DOCUMENT_NONFINITE", + lambda: gate.load_json_document(self.fixture("invalid/nonfinite.json")), + ) + + def test_unknown_and_missing_fields_fail_closed(self) -> None: + self.assert_reason( + "UNKNOWN_FIELD", + lambda: gate.validate_experiment(self.load("invalid/unknown-field.json")), + ) + self.assert_reason( + "MISSING_FIELD", + lambda: gate.validate_experiment(self.load("invalid/missing-field.json")), + ) + + def test_unrepresentable_numbers_and_impossible_timestamps_fail_closed( + self, + ) -> None: + huge_number = self.load("valid-finite/experiment.json") + physics = huge_number["physics"] + self.assertIsInstance(physics, dict) + model = physics["model"] + self.assertIsInstance(model, dict) + couplings = model["couplings"] + self.assertIsInstance(couplings, dict) + couplings["J"] = 10**4000 + self.assert_reason( + "VALUE_INVALID", + lambda: gate.validate_experiment(huge_number), + ) + + impossible_date = self.load("valid-finite/experiment.json") + provenance = impossible_date["provenance"] + self.assertIsInstance(provenance, dict) + provenance["created_at"] = "2026-02-30T12:00:00Z" + self.assert_reason( + "VALUE_INVALID", + lambda: gate.validate_experiment(impossible_date), + ) + + def test_unsupported_route_is_not_rewritten(self) -> None: + experiment = self.load("valid-finite/experiment.json") + experiment["capability"] = { + "capability_id": "quimb.finite_1d.dmrg", + "maturity": "catalogued", + "known_limitations": ["No promoted adapter binding."], + } + self.assert_reason( + "UNSUPPORTED_ROUTE", + lambda: gate.validate_experiment(experiment), + ) + + finite = self.load("valid-finite/experiment.json") + capability = finite["capability"] + self.assertIsInstance(capability, dict) + capability["known_limitations"] = ["none"] + self.assert_reason( + "UNSUPPORTED_ROUTE", + lambda: gate.validate_experiment(finite), + ) + + def test_bad_experiment_digest_and_binding_are_distinct(self) -> None: + finite = self.load("valid-finite/experiment.json") + bad_digest = self.load("valid-finite/evidence.json") + bad_digest["experiment_digest"] = "sha256:" + ("0" * 64) + self.assert_reason( + "EXPERIMENT_DIGEST_MISMATCH", + lambda: gate.evaluate( + finite, + bad_digest, + artifact_root=self.fixture("valid-finite/artifacts"), + ), + ) + + bad_binding = self.load("valid-finite/evidence.json") + binding = bad_binding["binding"] + self.assertIsInstance(binding, dict) + binding["adapter_id"] = "unapproved.adapter" + self.assert_reason( + "BINDING_MISMATCH", + lambda: gate.evaluate( + finite, + bad_binding, + artifact_root=self.fixture("valid-finite/artifacts"), + ), + ) + + def test_result_digest_is_checked(self) -> None: + finite = self.load("valid-finite/experiment.json") + evidence = self.load("valid-finite/evidence.json") + evidence["result_digest"] = "sha256:" + ("0" * 64) + self.assert_reason( + "RESULT_DIGEST_MISMATCH", + lambda: gate.evaluate( + finite, + evidence, + artifact_root=self.fixture("valid-finite/artifacts"), + ), + ) + + def test_artifact_digest_is_recomputed_from_raw_bytes(self) -> None: + finite = self.load("valid-finite/experiment.json") + evidence = self.load("valid-finite/evidence.json") + artifacts = evidence["artifacts"] + self.assertIsInstance(artifacts, list) + self.assertIsInstance(artifacts[0], dict) + artifacts[0]["digest"] = "sha256:" + ("0" * 64) + evidence["result_digest"] = gate.canonical_digest( + {key: value for key, value in evidence.items() if key != "result_digest"} + ) + self.assert_reason( + "ARTIFACT_DIGEST_MISMATCH", + lambda: gate.evaluate( + finite, + evidence, + artifact_root=self.fixture("valid-finite/artifacts"), + ), + ) + + def test_artifact_paths_must_be_canonical_relative_posix_paths(self) -> None: + for relative_path in ( + ".", + "\x00", + "artifacts/\x00", + "/backend-result.json", + "../backend-result.json", + "artifacts/../backend-result.json", + "./backend-result.json", + "artifacts/./backend-result.json", + "artifacts//backend-result.json", + "artifacts/", + r"artifacts\backend-result.json", + "C:backend-result.json", + ): + with self.subTest(relative_path=relative_path): + self.assert_reason( + "ARTIFACT_UNSAFE_PATH", + lambda relative_path=relative_path: gate._validate_relative_path( + relative_path, + "$.artifacts[0].relative_path", + ), + ) + + def test_artifact_path_errors_are_stable_json_without_tracebacks(self) -> None: + evidence = self.load("valid-finite/evidence.json") + artifacts = evidence["artifacts"] + self.assertIsInstance(artifacts, list) + self.assertIsInstance(artifacts[0], dict) + artifacts[0]["relative_path"] = "." + self.redigest_evidence(evidence) + with tempfile.TemporaryDirectory() as temporary: + evidence_path = Path(temporary) / "evidence.json" + evidence_path.write_text(json.dumps(evidence), encoding="utf-8") + completed = subprocess.run( + [ + sys.executable, + str(SOLUTION_ROOT / "gate.py"), + "evaluate", + str(self.fixture("valid-finite/experiment.json")), + str(evidence_path), + "--artifact-root", + str(self.fixture("valid-finite/artifacts")), + ], + check=False, + capture_output=True, + text=True, + ) + self.assertEqual(completed.returncode, 2) + self.assertEqual( + json.loads(completed.stdout)["reason_code"], + "ARTIFACT_UNSAFE_PATH", + ) + self.assertEqual(completed.stderr, "") + + def test_artifact_root_symlink_and_fifo_are_rejected_without_blocking(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + root = Path(temporary) + symlink = root / "artifact-root" + symlink.symlink_to( + self.fixture("valid-finite/artifacts"), + target_is_directory=True, + ) + self.assert_reason( + "ARTIFACT_UNSAFE_PATH", + lambda: gate.evaluate( + self.load("valid-finite/experiment.json"), + self.load("valid-finite/evidence.json"), + artifact_root=symlink, + ), + ) + + fifo_root = root / "fifo-root" + fifo_root.mkdir() + os.mkfifo(fifo_root / "fifo") + script = ( + "import pathlib,sys;" + f"sys.path.insert(0,{str(SOLUTION_ROOT)!r});" + "import gate;" + "gate._read_artifact_file(pathlib.Path(sys.argv[1]),'fifo',1024)" + ) + process = subprocess.Popen( + [sys.executable, "-c", script, str(fifo_root)], + stdout=subprocess.PIPE, + stderr=subprocess.PIPE, + text=True, + ) + try: + process.communicate(timeout=1) + except subprocess.TimeoutExpired: + process.kill() + process.communicate() + self.fail("FIFO artifact open blocked before file-type validation") + self.assertNotEqual(process.returncode, 0) + + def test_hardlinked_artifacts_are_rejected(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + artifacts = Path(temporary) / "artifacts" + shutil.copytree(self.fixture("valid-finite/artifacts"), artifacts) + os.link( + artifacts / "backend-request.json", + Path(temporary) / "backend-request-hardlink.json", + ) + self.assert_reason( + "ARTIFACT_UNSAFE_PATH", + lambda: gate.evaluate( + self.load("valid-finite/experiment.json"), + self.load("valid-finite/evidence.json"), + artifact_root=artifacts, + ), + ) + + def test_backend_result_artifact_must_contain_normalized_semantics(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + artifacts = Path(temporary) / "artifacts" + shutil.copytree(self.fixture("valid-finite/artifacts"), artifacts) + evidence = self.load("valid-finite/evidence.json") + self.replace_backend_bundle(artifacts, evidence, b"") + self.assert_reason( + "DOCUMENT_INVALID_JSON", + lambda: gate.evaluate( + self.load("valid-finite/experiment.json"), + evidence, + artifact_root=artifacts, + ), + ) + + def test_backend_bundle_json_is_strict_for_every_contract_boundary(self) -> None: + original = self.fixture( + "valid-finite/artifacts/backend-result.json" + ).read_bytes() + for mutation, expected in ( + ("duplicate", "DOCUMENT_DUPLICATE_KEY"), + ("nonfinite", "DOCUMENT_NONFINITE"), + ("unknown", "UNKNOWN_FIELD"), + ("missing", "MISSING_FIELD"), + ): + with ( + self.subTest(mutation=mutation), + tempfile.TemporaryDirectory() as temporary, + ): + artifacts = Path(temporary) / "artifacts" + shutil.copytree(self.fixture("valid-finite/artifacts"), artifacts) + evidence = self.load("valid-finite/evidence.json") + if mutation == "duplicate": + raw = original.replace( + b"{\n", + b'{\n "schema_version": "tn-agent.backend-result.v1",\n', + 1, + ) + elif mutation == "nonfinite": + raw = original.replace(b'"value": 1e-10', b'"value": NaN', 1) + else: + bundle = json.loads(original) + if mutation == "unknown": + bundle["unexpected"] = True + else: + del bundle["backend"] + bundle["result_digest"] = gate.canonical_digest( + { + key: value + for key, value in bundle.items() + if key != "result_digest" + } + ) + raw = (json.dumps(bundle, indent=2) + "\n").encode() + self.replace_backend_bundle(artifacts, evidence, raw) + self.assert_reason( + expected, + lambda evidence=evidence, artifacts=artifacts: gate.evaluate( + self.load("valid-finite/experiment.json"), + evidence, + artifact_root=artifacts, + ), + ) + + def test_raw_result_and_plan_id_must_match_the_normalized_bundle(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + artifacts = Path(temporary) / "artifacts" + shutil.copytree(self.fixture("valid-finite/artifacts"), artifacts) + evidence = self.load("valid-finite/evidence.json") + raw_path = artifacts / "backend-raw-result.json" + raw = raw_path.read_bytes() + b"\n" + raw_path.write_bytes(raw) + manifest = evidence["artifacts"] + self.assertIsInstance(manifest, list) + raw_artifact = next( + item + for item in manifest + if isinstance(item, dict) and item.get("role") == "backend_raw_result" + ) + raw_artifact["digest"] = "sha256:" + hashlib.sha256(raw).hexdigest() + raw_artifact["size_bytes"] = len(raw) + self.redigest_evidence(evidence) + self.assert_reason( + "RESULT_DIGEST_MISMATCH", + lambda: gate.evaluate( + self.load("valid-finite/experiment.json"), + evidence, + artifact_root=artifacts, + ), + ) + + evidence = self.load("valid-finite/evidence.json") + provenance = evidence["provenance"] + self.assertIsInstance(provenance, dict) + provenance["plan_id"] = "sha256:" + ("f" * 64) + self.redigest_evidence(evidence) + self.assert_reason( + "PROVENANCE_MISMATCH", + lambda: gate.evaluate( + self.load("valid-finite/experiment.json"), + evidence, + artifact_root=self.fixture("valid-finite/artifacts"), + ), + ) + + def test_raw_failed_status_and_forged_energy_cannot_reuse_normalized_pass( + self, + ) -> None: + for mutation, expected in ( + ("failed_status", "PROVENANCE_MISMATCH"), + ("energy_999", "RESULT_DIGEST_MISMATCH"), + ): + with ( + self.subTest(mutation=mutation), + tempfile.TemporaryDirectory() as temporary, + ): + artifacts = Path(temporary) / "artifacts" + shutil.copytree(self.fixture("valid-finite/artifacts"), artifacts) + evidence = self.load("valid-finite/evidence.json") + raw_result = json.loads( + (artifacts / "backend-raw-result.json").read_text() + ) + if mutation == "failed_status": + raw_result["status"] = "failed" + else: + raw_result["observables"]["energy"]["value"] = 999.0 + raw = (json.dumps(raw_result, indent=2) + "\n").encode() + self.replace_artifact( + artifacts, + evidence, + role="backend_raw_result", + raw=raw, + ) + self.assert_reason( + expected, + lambda evidence=evidence, artifacts=artifacts: gate.evaluate( + self.load("valid-finite/experiment.json"), + evidence, + artifact_root=artifacts, + ), + ) + + def test_reported_energy_drift_cannot_override_derived_delta(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + artifacts = Path(temporary) / "artifacts" + shutil.copytree(self.fixture("valid-finite/artifacts"), artifacts) + experiment = self.load("valid-finite/experiment.json") + evidence = self.load("valid-finite/evidence.json") + raw_result = json.loads((artifacts / "backend-raw-result.json").read_text()) + raw_result["convergence"][-1]["metrics"]["energy_drift"] = 999.0 + raw = (json.dumps(raw_result, indent=2) + "\n").encode() + self.replace_artifact( + artifacts, + evidence, + role="backend_raw_result", + raw=raw, + ) + summary = gate.validate_experiment(experiment) + binding = experiment["backend_binding"] + self.assertIsInstance(binding, dict) + plan_id = gate._expected_plan_id( + experiment, + str(summary["experiment_digest"]), + ) + request = gate._expected_request( + experiment, + str(summary["experiment_digest"]), + plan_id, + ) + manifest = evidence["artifacts"] + self.assertIsInstance(manifest, list) + raw_artifact = next( + item + for item in manifest + if isinstance(item, dict) and item.get("role") == "backend_raw_result" + ) + parsed_raw = gate._validate_raw_result( + raw, + request=request, + request_digest=gate.canonical_digest(request), + binding=binding, + route=gate.ROUTES[str(binding["capability_id"])], + ) + execution = evidence["execution"] + self.assertIsInstance(execution, dict) + reconstructed = gate._reconstruct_backend_bundle( + request_digest=gate.canonical_digest(request), + binding=binding, + raw=parsed_raw, + execution=execution, + raw_artifact=raw_artifact, + ) + self.write_backend_bundle(artifacts, evidence, reconstructed) + self.assert_reason( + "VALIDATOR_STATUS_INVALID", + lambda: gate.evaluate( + experiment, + evidence, + artifact_root=artifacts, + ), + ) + + def test_coherent_energy_forgery_fails_against_preregistered_reference( + self, + ) -> None: + with tempfile.TemporaryDirectory() as temporary: + artifacts = Path(temporary) / "artifacts" + shutil.copytree(self.fixture("valid-finite/artifacts"), artifacts) + experiment = self.load("valid-finite/experiment.json") + evidence = self.load("valid-finite/evidence.json") + raw_result = json.loads((artifacts / "backend-raw-result.json").read_text()) + raw_result["observables"]["energy"]["value"] = 999.0 + raw_result["convergence"][-2]["metrics"]["energy"] = 998.99999999 + raw_result["convergence"][-1]["metrics"]["energy"] = 999.0 + raw_result["convergence"][-1]["metrics"]["canonical_residual"] = 0.0 + raw_result["convergence"][-1]["metrics"]["symmetry_residual"] = 0.0 + self.coherently_rebuild_primary_evidence( + directory=artifacts, + experiment=experiment, + evidence=evidence, + raw_result=raw_result, + ) + self.assert_reason( + "VALIDATOR_THRESHOLD_FAILED", + lambda: gate.evaluate( + experiment, + evidence, + artifact_root=artifacts, + ), + ) + + def test_repeat_execution_and_raw_identity_must_be_distinct(self) -> None: + missing_repeat = self.load("valid-finite/evidence.json") + manifest = missing_repeat["artifacts"] + self.assertIsInstance(manifest, list) + manifest[:] = [ + item + for item in manifest + if not ( + isinstance(item, dict) and item.get("role") == "backend_repeat_stderr" + ) + ] + self.redigest_evidence(missing_repeat) + self.assert_reason( + "VALUE_INVALID", + lambda: gate.evaluate( + self.load("valid-finite/experiment.json"), + missing_repeat, + artifact_root=self.fixture("valid-finite/artifacts"), + ), + ) + + evidence = self.load("valid-finite/evidence.json") + evidence["repeat_execution"] = copy.deepcopy(evidence["execution"]) + self.redigest_evidence(evidence) + self.assert_reason( + "PROVENANCE_MISMATCH", + lambda: gate.evaluate( + self.load("valid-finite/experiment.json"), + evidence, + artifact_root=self.fixture("valid-finite/artifacts"), + ), + ) + + with tempfile.TemporaryDirectory() as temporary: + artifacts = Path(temporary) / "artifacts" + shutil.copytree(self.fixture("valid-finite/artifacts"), artifacts) + evidence = self.load("valid-finite/evidence.json") + primary_raw = (artifacts / "backend-raw-result.json").read_bytes() + self.replace_artifact( + artifacts, + evidence, + role="backend_repeat_raw_result", + raw=primary_raw, + ) + self.assert_reason( + "PROVENANCE_MISMATCH", + lambda: gate.evaluate( + self.load("valid-finite/experiment.json"), + evidence, + artifact_root=artifacts, + ), + ) + + def test_repeat_cosmetic_differences_remain_reported_only(self) -> None: + primary_bytes = self.fixture( + "valid-finite/artifacts/backend-raw-result.json" + ).read_bytes() + repeat_bytes = self.fixture( + "valid-finite/artifacts/backend-repeat-raw-result.json" + ).read_bytes() + cosmetic_warning = json.loads(repeat_bytes) + cosmetic_warning["warnings"] = ["cosmetic repeat warning"] + physics_equal = json.loads(primary_bytes) + physics_equal["warnings"] = ["physics-equal structural rebuild"] + cases = { + "trailing_whitespace": repeat_bytes + b" \n", + "cosmetic_warning": ( + json.dumps(cosmetic_warning, indent=2, allow_nan=False) + "\n" + ).encode(), + "physics_equal_full_rebuild": ( + json.dumps(physics_equal, indent=2, allow_nan=False) + "\n" + ).encode(), + } + for name, mutated_repeat in cases.items(): + with ( + self.subTest(case=name), + tempfile.TemporaryDirectory() as temporary, + ): + artifacts = Path(temporary) / "artifacts" + shutil.copytree(self.fixture("valid-finite/artifacts"), artifacts) + experiment = self.load("valid-finite/experiment.json") + evidence = self.load("valid-finite/evidence.json") + self.coherently_rebuild_repeat_evidence( + directory=artifacts, + experiment=experiment, + evidence=evidence, + repeat_raw_bytes=mutated_repeat, + ) + verdict = gate.evaluate( + experiment, + evidence, + artifact_root=artifacts, + ) + self.assertEqual(verdict["reason_code"], "ACCEPTANCE_PASSED") + reproduction = next( + result + for result in evidence["validator_results"] + if result["id"] == "reproducibility" + ) + self.assertEqual(reproduction["status"], "reported_only") + self.assertEqual(reproduction["reason_code"], "REPORTED_ONLY") + if name == "physics_equal_full_rebuild": + self.assertEqual(reproduction["metric_value"], 0.0) + + def test_reported_only_diagnostics_cannot_enter_required_acceptance(self) -> None: + experiment = self.load("valid-finite/experiment.json") + validators = experiment["validators"] + acceptance = experiment["acceptance"] + self.assertIsInstance(validators, list) + self.assertIsInstance(acceptance, dict) + variance = next( + item + for item in validators + if isinstance(item, dict) and item.get("id") == "variance" + ) + variance["policy"] = "required_pass" + variance["operator"] = "max" + variance["threshold"] = 1.0 + acceptance["reported_only_validator_ids"].remove("variance") + acceptance["required_validator_ids"].append("variance") + self.assert_reason( + "VALIDATOR_POLICY_MISMATCH", + lambda: gate.validate_experiment(experiment), + ) + + def test_validator_artifact_metric_is_recomputed_not_trusted(self) -> None: + with tempfile.TemporaryDirectory() as temporary: + artifacts = Path(temporary) / "artifacts" + shutil.copytree(self.fixture("valid-finite/artifacts"), artifacts) + evidence = self.load("valid-finite/evidence.json") + validator = json.loads((artifacts / "validator-evidence.json").read_text()) + convergence = next( + item for item in validator["results"] if item["id"] == "convergence" + ) + convergence["value"] = 0.0 + validator["result_digest"] = gate.canonical_digest( + { + key: value + for key, value in validator.items() + if key != "result_digest" + } + ) + raw = (json.dumps(validator, indent=2) + "\n").encode() + self.replace_artifact( + artifacts, + evidence, + role="validator_evidence", + raw=raw, + ) + self.assert_reason( + "VALIDATOR_STATUS_INVALID", + lambda: gate.evaluate( + self.load("valid-finite/experiment.json"), + evidence, + artifact_root=artifacts, + ), + ) + + def test_backend_bundle_is_the_source_of_observable_and_validator_semantics( + self, + ) -> None: + finite = self.load("valid-finite/experiment.json") + for mutation, expected in ( + ("observable_status", "OBSERVABLE_STATUS_INVALID"), + ("validator_metric", "VALIDATOR_STATUS_INVALID"), + ): + with ( + self.subTest(mutation=mutation), + tempfile.TemporaryDirectory() as temporary, + ): + artifacts = Path(temporary) / "artifacts" + shutil.copytree(self.fixture("valid-finite/artifacts"), artifacts) + evidence = self.load("valid-finite/evidence.json") + if mutation == "observable_status": + observable_results = evidence["observables"] + self.assertIsInstance(observable_results, list) + for item in observable_results: + self.assertIsInstance(item, dict) + item["status"] = "derived" + else: + validator_results = evidence["validator_results"] + self.assertIsInstance(validator_results, list) + convergence = next( + item + for item in validator_results + if isinstance(item, dict) and item.get("id") == "convergence" + ) + convergence["metric_value"] = 0.0 + self.redigest_evidence(evidence) + self.assert_reason( + expected, + lambda evidence=evidence, artifacts=artifacts: gate.evaluate( + finite, + evidence, + artifact_root=artifacts, + ), + ) + + def test_validator_metric_contracts_are_exact_and_nonnegative(self) -> None: + experiment = self.load("valid-finite/experiment.json") + validators = experiment["validators"] + self.assertIsInstance(validators, list) + convergence = next( + item + for item in validators + if isinstance(item, dict) and item.get("id") == "convergence" + ) + convergence["metric"] = "invented_metric" + self.assert_reason( + "VALIDATOR_POLICY_MISMATCH", + lambda: gate.validate_experiment(experiment), + ) + + evidence = self.load("valid-finite/evidence.json") + results = evidence["validator_results"] + self.assertIsInstance(results, list) + convergence_result = next( + item + for item in results + if isinstance(item, dict) and item.get("id") == "convergence" + ) + convergence_result["metric_value"] = -1.0 + self.redigest_evidence(evidence) + self.assert_reason( + "VALIDATOR_STATUS_INVALID", + lambda: gate.evaluate( + self.load("valid-finite/experiment.json"), + evidence, + artifact_root=self.fixture("valid-finite/artifacts"), + ), + ) + + def test_metric_cannot_claim_pass_above_preregistered_threshold(self) -> None: + finite = self.load("valid-finite/experiment.json") + validators = finite["validators"] + self.assertIsInstance(validators, list) + convergence = next( + item + for item in validators + if isinstance(item, dict) and item.get("id") == "convergence" + ) + self.assert_reason( + "VALIDATOR_THRESHOLD_FAILED", + lambda: gate._evaluate_threshold( + "convergence", + convergence, + {"metric_value": 1.0}, + ), + ) + + def test_library_is_append_only_and_cross_references_prior_records(self) -> None: + summary = gate.validate_library(SOLUTION_ROOT / "library" / "heuristics.jsonl") + self.assertEqual(summary["reason_code"], "OK") + self.assertGreaterEqual(summary["records"], 3) + + first = json.loads( + (SOLUTION_ROOT / "library" / "heuristics.jsonl") + .read_text(encoding="utf-8") + .splitlines()[0] + ) + first["revision"] = 2 + first["record_id"] = f"{first['heuristic_id']}@2" + with tempfile.TemporaryDirectory() as temporary: + bad_library = Path(temporary) / "heuristics.jsonl" + bad_library.write_text( + json.dumps(first, separators=(",", ":")) + "\n", + encoding="utf-8", + ) + self.assert_reason( + "LIBRARY_SEQUENCE_INVALID", + lambda: gate.validate_library(bad_library), + ) + + def test_library_revision_must_supersede_immediately_prior_record(self) -> None: + first = json.loads( + (SOLUTION_ROOT / "library" / "heuristics.jsonl") + .read_text(encoding="utf-8") + .splitlines()[0] + ) + second = json.loads(json.dumps(first)) + second["record_id"] = f"{first['heuristic_id']}@2" + second["revision"] = 2 + second["recorded_at"] = "2026-07-29T00:00:01Z" + second["supersedes"] = [first["record_id"]] + with tempfile.TemporaryDirectory() as temporary: + library = Path(temporary) / "heuristics.jsonl" + library.write_text( + "\n".join( + json.dumps(item, separators=(",", ":")) for item in (first, second) + ) + + "\n", + encoding="utf-8", + ) + self.assertEqual(gate.validate_library(library)["records"], 2) + second["supersedes"] = [] + library.write_text( + "\n".join( + json.dumps(item, separators=(",", ":")) for item in (first, second) + ) + + "\n", + encoding="utf-8", + ) + self.assert_reason( + "LIBRARY_SEQUENCE_INVALID", + lambda: gate.validate_library(library), + ) + + def test_library_revision_cannot_supersede_an_unrelated_heuristic(self) -> None: + records = [ + json.loads(line) + for line in (SOLUTION_ROOT / "library" / "heuristics.jsonl") + .read_text(encoding="utf-8") + .splitlines() + ] + revision = copy.deepcopy(records[0]) + revision["record_id"] = f"{revision['heuristic_id']}@2" + revision["revision"] = 2 + revision["recorded_at"] = "2026-07-29T00:03:00Z" + revision["supersedes"] = [records[0]["record_id"], records[1]["record_id"]] + with tempfile.TemporaryDirectory() as temporary: + library = Path(temporary) / "heuristics.jsonl" + library.write_text( + "\n".join( + json.dumps(item, separators=(",", ":")) + for item in (records[0], records[1], revision) + ) + + "\n", + encoding="utf-8", + ) + self.assert_reason( + "LIBRARY_SEQUENCE_INVALID", + lambda: gate.validate_library(library), + ) + + def test_library_grounding_rejects_tamper_traversal_and_missing_files( + self, + ) -> None: + original = json.loads( + (SOLUTION_ROOT / "library" / "heuristics.jsonl") + .read_text(encoding="utf-8") + .splitlines()[0] + ) + mutations = { + "source_tamper": ("source", "sha256", "sha256:" + ("0" * 64)), + "source_traversal": ("source", "uri", "../skills/method-mps/SKILL.md"), + "source_missing": ("source", "uri", "skills/missing/SKILL.md"), + "evidence_tamper": ("evidence", "sha256", "sha256:" + ("f" * 64)), + "evidence_traversal": ( + "evidence", + "uri", + "../skills/method-mps/SKILL.md", + ), + "evidence_missing": ("evidence", "uri", "skills/missing/SKILL.md"), + } + for name, (section, key, value) in mutations.items(): + with ( + self.subTest(case=name), + tempfile.TemporaryDirectory() as temporary, + ): + record = copy.deepcopy(original) + record[section][key] = value + library = Path(temporary) / "heuristics.jsonl" + library.write_text( + json.dumps(record, separators=(",", ":")) + "\n", + encoding="utf-8", + ) + self.assert_reason( + "LIBRARY_RECORD_INVALID", + lambda library=library: gate.validate_library(library), + ) + + def test_library_kind_and_path_must_match(self) -> None: + original = json.loads( + (SOLUTION_ROOT / "library" / "heuristics.jsonl") + .read_text(encoding="utf-8") + .splitlines()[0] + ) + mutations = ( + ("source", "repository_skill", "README.md", REPOSITORY_ROOT), + ("evidence", "method_card", "README.md", REPOSITORY_ROOT), + ("evidence", "workflow_card", "README.md", REPOSITORY_ROOT), + ("evidence", "contract_audit", "README.md", SOLUTION_ROOT), + ) + for section, kind, uri, root in mutations: + with ( + self.subTest(section=section, kind=kind), + tempfile.TemporaryDirectory() as temporary, + ): + record = copy.deepcopy(original) + grounded_file = root / uri + record[section]["kind"] = kind + record[section]["uri"] = uri + record[section]["sha256"] = ( + "sha256:" + hashlib.sha256(grounded_file.read_bytes()).hexdigest() + ) + library = Path(temporary) / "heuristics.jsonl" + library.write_text( + json.dumps(record, separators=(",", ":")) + "\n", + encoding="utf-8", + ) + self.assert_reason( + "LIBRARY_RECORD_INVALID", + lambda library=library: gate.validate_library(library), + ) + + def test_public_cli_emits_one_machine_readable_verdict(self) -> None: + command = [ + sys.executable, + str(SOLUTION_ROOT / "gate.py"), + "evaluate", + str(self.fixture("valid-finite/experiment.json")), + str(self.fixture("valid-finite/evidence.json")), + "--artifact-root", + str(self.fixture("valid-finite/artifacts")), + ] + completed = subprocess.run( + command, + check=False, + capture_output=True, + text=True, + ) + self.assertEqual(completed.returncode, 0, completed.stderr) + output = json.loads(completed.stdout) + self.assertEqual(output["reason_code"], "ACCEPTANCE_PASSED") + self.assertEqual(completed.stderr, "") + + def test_public_cli_usage_errors_are_also_machine_readable(self) -> None: + completed = subprocess.run( + [sys.executable, str(SOLUTION_ROOT / "gate.py"), "evaluate"], + check=False, + capture_output=True, + text=True, + ) + self.assertEqual(completed.returncode, 2) + self.assertEqual(json.loads(completed.stdout)["reason_code"], "CLI_USAGE_ERROR") + self.assertEqual(completed.stderr, "") + + def test_schema_files_are_json_and_fail_closed_at_every_object(self) -> None: + schema_paths = ( + SOLUTION_ROOT / "contracts" / "experiment-v1.schema.json", + SOLUTION_ROOT / "contracts" / "evidence-v1.schema.json", + SOLUTION_ROOT / "contracts" / "backend-result-v1.schema.json", + SOLUTION_ROOT / "contracts" / "validator-evidence-v1.schema.json", + SOLUTION_ROOT / "contracts" / "energy-reference-v1.schema.json", + SOLUTION_ROOT / "library" / "heuristic-v1.schema.json", + ) + for schema_path in schema_paths: + with self.subTest(schema=schema_path.name): + schema = json.loads(schema_path.read_text(encoding="utf-8")) + self.assertEqual( + schema["$schema"], "https://json-schema.org/draft/2020-12/schema" + ) + self.assertEqual(schema["additionalProperties"], False) + for definition in schema.get("$defs", {}).values(): + if ( + isinstance(definition, dict) + and definition.get("type") == "object" + ): + additional = definition.get("additionalProperties") + self.assertTrue( + additional is False or isinstance(additional, dict), + f"unbounded object definition: {definition}", + ) + + def test_public_json_schemas_accept_all_promoted_fixtures(self) -> None: + if importlib.util.find_spec("jsonschema") is None: + self.skipTest("optional jsonschema package is not installed") + from jsonschema import Draft202012Validator + + pairs: list[tuple[Path, Path]] = [] + for directory in ("valid-finite", "valid-infinite"): + for schema_name, fixture_name in ( + ("experiment-v1.schema.json", "experiment.json"), + ("evidence-v1.schema.json", "evidence.json"), + ("backend-result-v1.schema.json", "artifacts/backend-result.json"), + ( + "backend-result-v1.schema.json", + "artifacts/backend-repeat-result.json", + ), + ( + "validator-evidence-v1.schema.json", + "artifacts/validator-evidence.json", + ), + ( + "energy-reference-v1.schema.json", + "artifacts/energy-reference.json", + ), + ): + pairs.append( + ( + SOLUTION_ROOT / "contracts" / schema_name, + self.fixture(f"{directory}/{fixture_name}"), + ) + ) + for schema_path, fixture_path in pairs: + with self.subTest(fixture=fixture_path): + schema = json.loads(schema_path.read_text()) + fixture = json.loads(fixture_path.read_text()) + Draft202012Validator.check_schema(schema) + Draft202012Validator(schema).validate(fixture) + + experiment_schema = json.loads( + (SOLUTION_ROOT / "contracts" / "experiment-v1.schema.json").read_text() + ) + experiment_validator = Draft202012Validator(experiment_schema) + for directory in ("valid-finite", "valid-infinite"): + for dependency in ("energy", "variance"): + with self.subTest( + schema_route=directory, + missing_dependency=dependency, + ): + experiment = self.load(f"{directory}/experiment.json") + experiment["observables"].remove(dependency) + self.assertFalse(experiment_validator.is_valid(experiment)) + finite_with_sweeps = self.load("valid-finite/experiment.json") + finite_with_sweeps["numerics"]["min_sweeps"] = 1 + self.assertFalse(experiment_validator.is_valid(finite_with_sweeps)) + infinite_without_sweeps = self.load("valid-infinite/experiment.json") + del infinite_without_sweeps["numerics"]["min_sweeps"] + self.assertFalse(experiment_validator.is_valid(infinite_without_sweeps)) + infinite_nonunit_jxy = self.load("valid-infinite/experiment.json") + infinite_nonunit_jxy["physics"]["model"]["couplings"]["Jxy"] = 2.0 + self.assertFalse(experiment_validator.is_valid(infinite_nonunit_jxy)) + + library_schema = json.loads( + (SOLUTION_ROOT / "library" / "heuristic-v1.schema.json").read_text() + ) + Draft202012Validator.check_schema(library_schema) + library_validator = Draft202012Validator(library_schema) + for line in ( + (SOLUTION_ROOT / "library" / "heuristics.jsonl").read_text().splitlines() + ): + library_validator.validate(json.loads(line)) + library_record = json.loads( + (SOLUTION_ROOT / "library" / "heuristics.jsonl") + .read_text(encoding="utf-8") + .splitlines()[0] + ) + for section, kind in ( + ("source", "repository_skill"), + ("evidence", "method_card"), + ("evidence", "workflow_card"), + ("evidence", "contract_audit"), + ): + with self.subTest(schema_section=section, schema_kind=kind): + mutation = copy.deepcopy(library_record) + mutation[section]["kind"] = kind + mutation[section]["uri"] = "README.md" + self.assertFalse(library_validator.is_valid(mutation)) + + def test_public_calibration_evidence_is_digest_bound(self) -> None: + calibration_root = SOLUTION_ROOT / "calibration" + candidate_path = calibration_root / "blind-candidates.json" + controls_path = calibration_root / "weak-controls.json" + self.assertEqual( + hashlib.sha256(candidate_path.read_bytes()).hexdigest(), + "101bf35d22607b08ad0b160b893392e021bf7f8ac9d54aaefce91007f3b37be8", + ) + + candidates = json.loads(candidate_path.read_text(encoding="utf-8")) + self.assertIsInstance(candidates, list) + self.assertEqual(len(candidates), 5) + + report = json.loads( + (calibration_root / "report.json").read_text(encoding="utf-8") + ) + manifest = json.loads( + (calibration_root / "manifest.json").read_text(encoding="utf-8") + ) + controls = json.loads(controls_path.read_text(encoding="utf-8")) + self.assertEqual( + "sha256:" + hashlib.sha256(controls_path.read_bytes()).hexdigest(), + manifest["weak_controls_digest"], + ) + self.assertEqual( + report["candidate_digests"], + [gate.canonical_digest(candidate) for candidate in candidates], + ) + for document in (manifest, report): + self.assertEqual( + document["digest"], + gate.canonical_digest( + {key: value for key, value in document.items() if key != "digest"} + ), + ) + + self.assertTrue(report["passed"]) + self.assertEqual(report["manifest_digest"], manifest["digest"]) + self.assertEqual( + report["sealed_target_digests"], + [target["statement_digest"] for target in manifest["targets"]], + ) + + candidate_count = len(candidates) + self.assertEqual( + report["scores"]["meaningful_gap_recovery"], + sum(bool(candidate["meaningful_gap"]) for candidate in candidates) + / candidate_count, + ) + self.assertEqual( + report["scores"]["executable_gate"], + sum( + bool(candidate["gate_executable"]) + and bool(candidate["gate_attack_passed"]) + for candidate in candidates + ) + / candidate_count, + ) + candidate_strength = ( + sum( + ( + float(candidate["novelty_score"]) + + float(candidate["publishability_score"]) + ) + / 2 + for candidate in candidates + ) + / candidate_count + ) + control_strength = sum( + (float(control["novelty_score"]) + float(control["publishability_score"])) + / 2 + for control in controls + ) / len(controls) + self.assertEqual( + report["scores"]["strong_weak_separation"], + candidate_strength - control_strength, + ) + + available = set(range(candidate_count)) + matched = 0 + minimum = float(manifest["thresholds"]["minimum_match_similarity"]) + for target in sorted(manifest["targets"], key=lambda item: item["target_id"]): + target_literature = set(target["literature_ids"]) + ranked = sorted( + ( + ( + len( + target_literature.intersection( + candidates[index]["literature_ids"] + ) + ) + / len( + target_literature.union(candidates[index]["literature_ids"]) + ), + index, + ) + for index in available + ), + key=lambda item: (-item[0], candidates[item[1]]["candidate_id"]), + ) + similarity, candidate_index = ranked[0] + if similarity >= minimum: + available.remove(candidate_index) + matched += 1 + self.assertEqual(matched, 4) + self.assertEqual( + report["scores"]["hidden_target_match"], + matched / len(manifest["targets"]), + ) + + def test_public_calibration_excludes_sealed_payloads_and_states_limits( + self, + ) -> None: + calibration_root = SOLUTION_ROOT / "calibration" + self.assertEqual( + { + path.relative_to(calibration_root).as_posix() + for path in calibration_root.rglob("*") + if path.is_file() + }, + { + ".gitignore", + "README.md", + "blind-candidates.json", + "manifest.json", + "report.json", + "weak-controls.json", + }, + ) + ignore_rules = ( + (calibration_root / ".gitignore").read_text(encoding="utf-8").splitlines() + ) + self.assertIn("sealed/*", ignore_rules) + self.assertIn("!sealed/README.md", ignore_rules) + documents = "\n".join( + ( + (calibration_root / "README.md").read_text(encoding="utf-8"), + (SOLUTION_ROOT / "README.md").read_text(encoding="utf-8"), + ) + ) + self.assertIn("4 of 5", documents) + self.assertIn("self-reported", documents) + self.assertIn("do not independently prove", documents) + + def test_fixture_provenance_digests_bind_repository_sources(self) -> None: + experiment = self.load("valid-finite/experiment.json") + provenance = experiment["provenance"] + self.assertIsInstance(provenance, dict) + sources = provenance["sources"] + self.assertIsInstance(sources, list) + for source in sources: + self.assertIsInstance(source, dict) + source_path = REPOSITORY_ROOT / source["uri"] + observed = "sha256:" + hashlib.sha256(source_path.read_bytes()).hexdigest() + self.assertEqual(observed, source["sha256"]) + + def test_reason_code_document_covers_the_executable_registry(self) -> None: + documentation = (SOLUTION_ROOT / "contracts" / "reason-codes.md").read_text( + encoding="utf-8" + ) + for reason_code in gate.GATE_REASON_CODES: + with self.subTest(reason_code=reason_code): + self.assertIn(f"`{reason_code}`", documentation) + + +if __name__ == "__main__": + unittest.main()