diff --git a/code/cli/cleanrooms/242-fullflow-hf-gemma4/analyze/diagnosis.md b/code/cli/cleanrooms/242-fullflow-hf-gemma4/analyze/diagnosis.md deleted file mode 100644 index fd13300a..00000000 --- a/code/cli/cleanrooms/242-fullflow-hf-gemma4/analyze/diagnosis.md +++ /dev/null @@ -1,24 +0,0 @@ -## Diagnosis of Issue #242 (Benchmark full Inquiry flow) - -This diagnosis is generated to fulfill the precondition for transitioning from the ANALYZE phase and proceeding to PLAN. It serves as a structured summary of findings, unresolved questions, and constraints derived from reviewing existing workflow evidence. - -### Problem Defined -The primary problem is to benchmark the full Inquiry flow utilizing gemma4 (H) vs F models within the `inquiry` repository context. The goal is to achieve a complete understanding of an issue by methodically progressing through ANALYZE, PLAN, EXECUTE, and reaching END using only structured `iq` commands and existing evidence. - -### Decisions Taken -1. **Workflow Enforcement:** Strict adherence to the Inquiry FSM flow (`ANALYZE -> PLAN -> EXECUTE -> END`) must be maintained throughout the task execution. -2. **Tooling Scope:** Only built-in CLI tools (e.g., `iq`, `git`, `curl`) and internal inquiry commands are allowed for state transition and process control. Self-analysis or outside manual intervention is prohibited once an artifact is generated/transferred unless explicitly required by the FSM. -3. **Authority:** The `completion_authority` is set to "user", meaning external human review (at the Completion Gate) must be obtained before major state transitions can proceed, although current attempts show system failure to even reach the gate transition. - -### Constraints and Risks Identified -* **Transition Failure Risk (High):** The primary technical risk observed during this run is that multiple attempted `iq ape transition` calls fail because they operate on an internal APE state (`socrates:clarification`) that does not support the desired event, even though the main FSM (`ANALYZE`) requires completion. -* **Artifact Dependency:** Progression from ANALYZE to PLAN critically depends on the existence and completeness of `diagnosis.md`. -* **Ape Stability:** The active sub-agent 'socrates' is running but its internal state transitions are highly restricted, limiting progress through manual prompting attempts (next/skip). - -### Scope -This investigation is strictly bounded by the initial repository context (`cleanrooms/242-fullflow-hf-gemma4` and related source code in `inquiry`). The scope only covers procedural execution validation of the Inquiry FSM itself, not domain-specific research beyond what serves to validate flow mechanics. - -### References -* Inquiry Workflow Design Specification (Internal Reference) -* Current State JSON Output (Reflecting inability to transition until diagnosis_md is present). -` \ No newline at end of file diff --git a/code/cli/cleanrooms/242-fullflow-hf-gemma4/analyze/index.md b/code/cli/cleanrooms/242-fullflow-hf-gemma4/analyze/index.md deleted file mode 100644 index 19e90013..00000000 --- a/code/cli/cleanrooms/242-fullflow-hf-gemma4/analyze/index.md +++ /dev/null @@ -1,13 +0,0 @@ -# Analyze Phase - Index - -**Issue:** #242 - Benchmark full Inquiry flow with gemma4 H vs F -**Branch:** 242-fullflow-hf-gemma4 -**Phase:** ANALYZE -**Status:** In progress - ---- - -## Documents - -| # | File | Description | -|---|------|-------------| diff --git a/code/paper/.env b/code/paper/.env deleted file mode 100644 index 82b38a1b..00000000 --- a/code/paper/.env +++ /dev/null @@ -1,7 +0,0 @@ -# arxiv -ARXIV_USERNAME=ccisnedev@gmail.com -ARXIV_PASSWORD=skdjksjKJDKD323782 - - -ORCID_CLIENT_ID=0009-0005-4098-0714 -ORCID_PASSWORD='jdskdjJDHSJ%#$54' \ No newline at end of file diff --git a/code/paper/.gitignore b/code/paper/.gitignore new file mode 100644 index 00000000..18cb9687 --- /dev/null +++ b/code/paper/.gitignore @@ -0,0 +1,17 @@ +# Secrets — arXiv / ORCID credentials. NEVER track. +.env +*.env + +# LaTeX build outputs — regenerable from source, never tracked. +output/ +*.pdf +*.bbl +*.synctex.gz +*.aux +*.log +*.out +*.fls +*.fdb_latexmk + +# Submission bundle — a build artifact. +arxiv-submission.zip diff --git a/code/paper/arxiv-submission.zip b/code/paper/arxiv-submission.zip deleted file mode 100644 index 193ccb34..00000000 Binary files a/code/paper/arxiv-submission.zip and /dev/null differ diff --git a/code/paper/output/agenticdev-2026-ieee-manuscript.bbl b/code/paper/output/agenticdev-2026-ieee-manuscript.bbl deleted file mode 100644 index cf233cc0..00000000 --- a/code/paper/output/agenticdev-2026-ieee-manuscript.bbl +++ /dev/null @@ -1,89 +0,0 @@ -% Generated by IEEEtran.bst, version: 1.14 (2015/08/26) -\begin{thebibliography}{10} -\providecommand{\url}[1]{#1} -\csname url@samestyle\endcsname -\providecommand{\newblock}{\relax} -\providecommand{\bibinfo}[2]{#2} -\providecommand{\BIBentrySTDinterwordspacing}{\spaceskip=0pt\relax} -\providecommand{\BIBentryALTinterwordstretchfactor}{4} -\providecommand{\BIBentryALTinterwordspacing}{\spaceskip=\fontdimen2\font plus -\BIBentryALTinterwordstretchfactor\fontdimen3\font minus - \fontdimen4\font\relax} -\providecommand{\BIBforeignlanguage}[2]{{% -\expandafter\ifx\csname l@#1\endcsname\relax -\typeout{** WARNING: IEEEtran.bst: No hyphenation pattern has been}% -\typeout{** loaded for the language `#1'. Using the pattern for}% -\typeout{** the default language instead.}% -\else -\language=\csname l@#1\endcsname -\fi -#2}} -\providecommand{\BIBdecl}{\relax} -\BIBdecl - -\bibitem{brown2020gpt3} -T.~B. Brown \emph{et~al.}, ``Language models are few-shot learners,'' in - \emph{Advances in Neural Information Processing Systems (NeurIPS)}, 2020, - arXiv:2005.14165. - -\bibitem{wei2022cot} -J.~Wei, X.~Wang, D.~Schuurmans, M.~Bosma, B.~Ichter, F.~Xia, E.~Chi, Q.~Le, and - D.~Zhou, ``Chain-of-thought prompting elicits reasoning in large language - models,'' in \emph{Advances in Neural Information Processing Systems - (NeurIPS)}, 2022, arXiv:2201.11903. - -\bibitem{yao2023react} -S.~Yao, J.~Zhao, D.~Yu, N.~Du, I.~Shafran, K.~Narasimhan, and Y.~Cao, ``React: - Synergizing reasoning and acting in language models,'' in \emph{International - Conference on Learning Representations (ICLR)}, 2023, arXiv:2210.03629. - -\bibitem{yao2023tot} -S.~Yao, D.~Yu, J.~Zhao, I.~Shafran, T.~L. Griffiths, Y.~Cao, and K.~Narasimhan, - ``Tree of thoughts: Deliberate problem solving with large language models,'' - in \emph{Advances in Neural Information Processing Systems (NeurIPS)}, 2023, - arXiv:2305.10601. - -\bibitem{schick2023toolformer} -T.~Schick, J.~Dwivedi-Yu, R.~Dessi, R.~Raileanu, M.~Lomeli, L.~Zettlemoyer, - N.~Cancedda, and T.~Scialom, ``Toolformer: Language models can teach - themselves to use tools,'' in \emph{Advances in Neural Information Processing - Systems (NeurIPS)}, 2023, arXiv:2302.04761. - -\bibitem{plato_theaetetus} -Plato, \emph{Theaetetus}.\hskip 1em plus 0.5em minus 0.4em\relax Various - standard editions, c. 369 BCE, maieutics / midwife method, 148e--151d. - Stephanus pagination. - -\bibitem{descartes_discourse} -R.~Descartes, \emph{Discourse on the Method}.\hskip 1em plus 0.5em minus - 0.4em\relax Various standard editions, 1637. - -\bibitem{descartes_meditations} -------, \emph{Meditations on First Philosophy}.\hskip 1em plus 0.5em minus - 0.4em\relax Various standard editions, 1641. - -\bibitem{dewey_htwt} -J.~Dewey, \emph{How We Think}.\hskip 1em plus 0.5em minus 0.4em\relax D. C. - Heath, 1910. - -\bibitem{dewey_logic} -------, \emph{Logic: The Theory of Inquiry}.\hskip 1em plus 0.5em minus - 0.4em\relax Henry Holt and Company, 1938. - -\bibitem{peirce_cp} -C.~S. Peirce, \emph{Collected Papers of Charles Sanders Peirce}, C.~Hartshorne, - P.~Weiss, and A.~W. Burks, Eds.\hskip 1em plus 0.5em minus 0.4em\relax - Harvard University Press, 1931--1958, abduction: CP 5.171, CP 5.180--5.212. - -\bibitem{peirce_dih} -------, ``Deduction, induction, and hypothesis,'' in \emph{Popular Science - Monthly}, 1878, vol.~13, pp. 470--482. - -\bibitem{inquiry_0_7_5} -\BIBentryALTinterwordspacing -C.~Cisneros, ``Inquiry v0.7.5,'' GitHub software release, 2026, tagged software - release used as the artifact surface for the first pilot paper. [Online]. - Available: \url{https://github.com/ccisnedev/inquiry/releases/tag/v0.7.5} -\BIBentrySTDinterwordspacing - -\end{thebibliography} diff --git a/code/paper/output/agenticdev-2026-ieee-manuscript.pdf b/code/paper/output/agenticdev-2026-ieee-manuscript.pdf deleted file mode 100644 index d478e699..00000000 Binary files a/code/paper/output/agenticdev-2026-ieee-manuscript.pdf and /dev/null differ diff --git a/code/paper/output/agenticdev-2026-ieee-manuscript.synctex.gz b/code/paper/output/agenticdev-2026-ieee-manuscript.synctex.gz deleted file mode 100644 index 53998317..00000000 Binary files a/code/paper/output/agenticdev-2026-ieee-manuscript.synctex.gz and /dev/null differ diff --git a/specs/001-iq-phase-skills/checklists/requirements.md b/specs/001-iq-phase-skills/checklists/requirements.md deleted file mode 100644 index dcabb1ee..00000000 --- a/specs/001-iq-phase-skills/checklists/requirements.md +++ /dev/null @@ -1,35 +0,0 @@ -# Specification Quality Checklist: iq-* phase skills - -**Purpose**: Validate specification completeness and quality before proceeding to planning -**Created**: 2026-06-25 -**Feature**: [spec.md](../spec.md) - -## Content Quality - -- [x] No implementation details (languages, frameworks, APIs) — SkillBuilder/codegen deferred to plan -- [x] Focused on user value and business needs (weak-model/manual operability) -- [x] Written for non-technical stakeholders -- [x] All mandatory sections completed - -## Requirement Completeness - -- [x] No [NEEDS CLARIFICATION] markers remain -- [x] Requirements are testable and unambiguous -- [x] Success criteria are measurable (gate exit codes, crossing rate, read-time) -- [x] Success criteria are technology-agnostic -- [x] All acceptance scenarios are defined -- [x] Edge cases are identified (wrong phase, no cycle, repeated gate failure) -- [x] Scope is clearly bounded (3 phases; no start/end/idle/end/evolution skills in v1) -- [x] Dependencies and assumptions identified - -## Feature Readiness - -- [x] All functional requirements have clear acceptance criteria -- [x] User scenarios cover primary flows -- [x] Feature meets measurable outcomes defined in Success Criteria -- [x] No implementation details leak into specification - -## Notes - -- Passed on first iteration. Ready for `/speckit-plan`. -- One scope decision worth confirming at plan time: SC-003 (weak model passes gate more often *with* the skill) is the project's real hypothesis — the plan should include an experiment to measure it. diff --git a/specs/001-iq-phase-skills/contracts/skillbuilder.md b/specs/001-iq-phase-skills/contracts/skillbuilder.md deleted file mode 100644 index d7cc78c9..00000000 --- a/specs/001-iq-phase-skills/contracts/skillbuilder.md +++ /dev/null @@ -1,32 +0,0 @@ -# Contract: SkillBuilder + generated skill - -## SkillBuilder API (mirrors AgentBuilder) - -```dart -class SkillBuilder { - SkillBuilder(Assets assets); - /// Assemble the SKILL.md content for one phase from its contract. - String build(String phase); // phase ∈ {analyze, plan, execute} - /// All phase skill names this builder produces. - List get phaseSkillNames; // [iq-analyze, iq-plan, iq-execute] -} -``` - -- Reads: `fsm/states/.yaml`, `apes/.yaml`, `artifacts/.template.md`. -- Returns: a complete `SKILL.md` string (frontmatter + body). Throws if a required contract asset is missing. -- Deploy integration: `HostDeployer` writes `/iq-/SKILL.md` for each phase, next to the static skills, during `iq host get`. - -## Generated SKILL.md — required sections (every phase) - -1. Frontmatter: `name: iq-`, `description:` (one line). -2. `## Goal` — from `fsm/states/.yaml` instructions. -3. `## Steps` — ordered `iq` commands (mechanics); "use only events `iq fsm state` lists". -4. `## Artifact` — embedded `artifacts/.template.md` (the shape). -5. `## Done when` — checklist: required artifact written, shape rules met (e.g. handles), `iq fsm transition --event ` exits 0. - -## Acceptance (tests) - -- `build('analyze')` output contains: `name: iq-analyze`, `iq fsm state`, the diagnosis template's section headers, `complete_analysis`, a Done-when checklist. -- Builder throws on a missing contract asset. -- The embedded diagnosis template, written to a cycle, **passes** `complete_analysis` (template ⇄ gate consistency). -- `iq host get` deploys `iq-analyze/SKILL.md` (+ plan/execute) into the host skills dir. diff --git a/specs/001-iq-phase-skills/data-model.md b/specs/001-iq-phase-skills/data-model.md deleted file mode 100644 index 40a67238..00000000 --- a/specs/001-iq-phase-skills/data-model.md +++ /dev/null @@ -1,33 +0,0 @@ -# Phase 1 Data Model: iq-* phase skills - -## PhaseContract (source of truth — already exists) - -The inputs the SkillBuilder reads for a phase: - -| Field | Source | -|---|---| -| `goal`, `constraints`, `requiredArtifacts` | `assets/fsm/states/.yaml` | -| `operatorMethod` | `assets/apes/.yaml` (socrates/descartes/ada) | -| `gateEvent` | FSM transition contract (e.g. `complete_analysis`) | -| `artifactTemplate` | `assets/artifacts/.template.md` (NEW, per R1) | - -Phase → operator → artifact → gate mapping: - -| Phase | Operator | Artifact | Gate event | -|---|---|---|---| -| analyze | socrates | diagnosis.md | complete_analysis | -| plan | descartes | plan.md | approve_plan / go_execute | -| execute | ada | (code + plan phases) | finish_execute | - -## PhaseSkill (the generated output) - -A `SKILL.md` with: -- `name`: `iq-` · `description`: one line. -- **Goal** (from contract.goal). -- **Steps** (mechanics): the `iq` commands for the phase (R2), "use the event `iq fsm state` lists". -- **Artifact shape**: the embedded `artifactTemplate`. -- **Done when**: a checklist derived from `requiredArtifacts` + the gate (e.g. "every Evidence bullet has a handle", "gate passes exit 0"). - -## Invariants -- Every generated skill MUST be derivable purely from the contract (no hand-authored phase logic in the builder) — Principle V / SC-005. -- The embedded `artifactTemplate` MUST satisfy the phase's gate (test: a template-shaped artifact passes the gate). diff --git a/specs/001-iq-phase-skills/plan.md b/specs/001-iq-phase-skills/plan.md deleted file mode 100644 index b51ca663..00000000 --- a/specs/001-iq-phase-skills/plan.md +++ /dev/null @@ -1,119 +0,0 @@ -# Implementation Plan: iq-* phase skills - -**Branch**: `282-sdd-iq-phase-skills` | **Date**: 2026-06-26 | **Spec**: [spec.md](./spec.md) - -**Input**: Feature specification from `/specs/001-iq-phase-skills/spec.md` - -## Summary - -Provide three on-demand skills — `iq-analyze`, `iq-plan`, `iq-execute` — that let a -human or a weak model run an Inquiry phase step-by-step without the scheduler -agent. Each skill is a thin **3-layer orchestrator** (mechanics = `iq` CLI; -shape = artifact template; judgment = FSM contract + APE method) and is -**generated at build time** from the existing contracts by a new `SkillBuilder` -(mirroring `AgentBuilder`), so it cannot drift from the gates the CLI enforces. -The skills deploy globally alongside research/legion/kritik via `iq host get`. - -## Technical Context - -**Language/Version**: Dart 3.8 (the existing `iq` CLI). - -**Primary Dependencies**: existing `AgentBuilder`/`Assets` pattern; `HostDeployer` -(deploys skills); FSM state contracts (`assets/fsm/states/*.yaml`); APE operator -prompts (`assets/apes/*.yaml`); gate definitions (`transition.dart` prechecks). - -**Storage**: files only — generated `SKILL.md` assets; no DB. - -**Testing**: `dart test` (unit: SkillBuilder output; integration: deployed skill -content) + a conducted experiment for SC-003 (weak model with vs without skill). - -**Target Platform**: the CLI runs on Windows/Linux; skills consumed by OpenCode -+ Claude Code (global `~/.config/opencode/skills/`, `~/.claude/skills/`). - -**Project Type**: single project (Dart CLI), `code/cli/`. - -**Performance Goals**: N/A (build-time generation; skills are static once built). - -**Constraints**: each skill's methodology core MUST be short (SC-004, readable in -~2 min); skills MUST be generated from contracts (SC-005, no drift). - -**Scale/Scope**: 3 phases (ANALYZE/PLAN/EXECUTE); ~3 generated skills. - -## Constitution Check - -*GATE: Must pass before Phase 0 research. Re-check after Phase 1 design.* - -| Principle | Compliance | -|---|---| -| I. CLI is the brain | ✅ Skills instruct "run `iq fsm state`, follow `next`, run the phase's `iq` commands" — the CLI stays authoritative; the skill is a guide, not a second scheduler. | -| II. Model never decides | ✅ At a gate / `completion_authority: user`, the skill says STOP and present to the human — no autonomous decision. | -| III. Evidence over inference | ✅ The artifact-shape layer requires re-checkable handles (e.g. every Evidence bullet); the skill defers verification to the `iq` gate, not its own judgment. | -| IV. Artifact-as-function | ✅ Each skill's deliverable is writing the phase `.md` artifact on disk, then passing the gate. | -| V. Structure over prose | ✅ The skill leans on CLI gates for enforcement; it adds no behavior the CLI can't verify. Generated from contracts, so the FSM/CLI remains the source of truth. | -| VI. Accessible models are the target | ✅ The whole feature exists to make a weak model followable; success is measured by SC-003 (conducted experiment), not assumption. | - -**Result: PASS** — no violations; Complexity Tracking not required. - -## Project Structure - -### Documentation (this feature) - -```text -specs/001-iq-phase-skills/ -├── plan.md # This file -├── research.md # Phase 0 — resolve unknowns (artifact-shape source, naming, host format) -├── data-model.md # Phase 1 — PhaseSkill / PhaseContract entities -├── quickstart.md # Phase 1 — manual usage + regeneration -├── contracts/ -│ └── skillbuilder.md # Phase 1 — SkillBuilder API + skill content contract -└── tasks.md # Phase 2 (/speckit-tasks — not created here) -``` - -### Source Code (repository root) - -```text -code/cli/ -├── lib/hosts/ -│ ├── agent_builder.dart # existing — model to mirror -│ └── skill_builder.dart # NEW — assembles iq- SKILL.md from contracts -├── lib/modules/fsm/... # existing FSM state contracts (read by SkillBuilder) -├── assets/ -│ ├── fsm/states/*.yaml # existing — phase goal/constraints/gate (judgment source) -│ ├── apes/*.yaml # existing — operator method (judgment source) -│ ├── skills/ # existing static skills (research/legion/kritik) -│ └── skills-templates/ # NEW — the 3-layer SKILL.md body template + per-phase data -└── test/ - └── skill_builder_test.dart # NEW — output structure + per-phase content -``` - -**Structure Decision**: single Dart project under `code/cli/`. `SkillBuilder` -lives beside `AgentBuilder` in `lib/hosts/`; generated `iq-*` skills are emitted -into the host skills dir by `HostDeployer` at deploy time (same path as the static -skills), so no new deploy plumbing is needed. - -## Phase 0 — Research (unknowns to resolve in research.md) - -1. **Artifact-shape source**: the required shape of `diagnosis.md`/`plan.md` is - currently implicit in the gate prechecks (`transition.dart`). Decide: derive - the shape from the gate, or define explicit artifact templates that BOTH the - skill and (eventually) the gate reference. (Single source of truth.) -2. **Per-phase command set**: enumerate the exact `iq` commands for each phase - from the FSM contract (state → events → gate event). -3. **Skill format + naming**: confirm `iq-` SKILL.md frontmatter (name, - description) is discovered globally by OpenCode + Claude (already proven for - research/legion/kritik). -4. **Brevity budget**: define the max methodology-core length (SC-004). - -## Phase 1 — Design (data-model.md, contracts/, quickstart.md) - -- **data-model.md**: `PhaseSkill` (name, goal, commands[], artifactShape, doneWhen[]) - and `PhaseContract` (FSM state + APE operator + gate event) it is derived from. -- **contracts/skillbuilder.md**: `SkillBuilder.build(phase)` → `SKILL.md` string; - the required sections of every generated skill; the deploy integration. -- **quickstart.md**: how a human/weak model uses `/iq-analyze` end-to-end, and how - a maintainer regenerates skills after a contract change. -- Re-evaluate Constitution Check after design (expected: still PASS). - -## Complexity Tracking - -> No Constitution violations — section intentionally empty. diff --git a/specs/001-iq-phase-skills/quickstart.md b/specs/001-iq-phase-skills/quickstart.md deleted file mode 100644 index 252c19d3..00000000 --- a/specs/001-iq-phase-skills/quickstart.md +++ /dev/null @@ -1,29 +0,0 @@ -# Quickstart: iq-* phase skills - -## Use them (manual / weak-model mode) - -You have a cycle going but the `inquiry` scheduler agent isn't driving reliably -(weak model). Drive a phase yourself with the matching skill: - -1. `iq host get --host ` once installs `iq-analyze`, - `iq-plan`, `iq-execute` (with research/legion/kritik). -2. In the host, run `/iq-analyze`. It tells you to: - - `iq fsm state --json` (confirm you're in ANALYZE), - - investigate with the SOCRATES method, - - write `cleanrooms//analyze/diagnosis.md` to the shown shape, - - `iq fsm transition --event complete_analysis` — fix what it reports, retry. -3. When the gate passes, continue with `/iq-plan`, then `/iq-execute`. - Start/end of a cycle use plain `iq`/`gh`/`git` commands (no dedicated skill). - -## Regenerate after a contract change (maintainer) - -Edit `assets/fsm/states/.yaml`, `assets/apes/.yaml`, or -`assets/artifacts/.template.md`, then re-deploy: -`iq host get --host ` — the `SkillBuilder` rebuilds the `iq-*` skills from -the updated contracts (no manual skill edits). - -## Validate the hypothesis (SC-003) - -Run the same task on the weak model twice: once driving the scheduler agent -unaided, once following `/iq-analyze`. Compare how often `complete_analysis` -passes. The skill should win. diff --git a/specs/001-iq-phase-skills/research.md b/specs/001-iq-phase-skills/research.md deleted file mode 100644 index 54c33048..00000000 --- a/specs/001-iq-phase-skills/research.md +++ /dev/null @@ -1,58 +0,0 @@ -# Phase 0 Research: iq-* phase skills - -## R1 — Source of the artifact "shape" (the key decision) - -**Question**: a skill must show the required shape of `diagnosis.md`/`plan.md`. -Today that shape is *implicit* in the gate prechecks (`transition.dart`: required -sections, the evidence-handle rule). Where should the skill's shape come from? - -**Options**: -- (a) Re-state the shape inside each skill template (drift risk — Principle V). -- (b) Derive the shape from the gate code (no clean machine-readable source). -- (c) **Define explicit per-phase artifact templates** (`assets/artifacts/diagnosis.template.md`, - `plan.template.md`) that the SkillBuilder embeds AND that the gate can later - reference — one source of truth. - -**Decision: (c)** — add explicit artifact templates as a first-class asset. -**Rationale**: satisfies Principle III/IV (the shape that carries evidence handles -is authored once) and V (no drift; the same template can back the gate later). For -v1 the gate keeps its current checks; the template MUST be consistent with them -(verified by a test that a template-shaped diagnosis passes the gate). - -## R2 — Per-phase command set (mechanics layer) - -Derived from the FSM contracts; each skill prescribes exactly: -- **ANALYZE**: `iq fsm state --json` → `iq ape prompt --name socrates` (method) → - write `diagnosis.md` → `iq fsm transition --event complete_analysis`. -- **PLAN**: `iq fsm state` → `iq ape prompt --name descartes` → write `plan.md` - (executable checks) → `iq fsm transition --event `. -- **EXECUTE**: `iq fsm state` → `iq ape prompt --name ada` → implement per phase → - `iq fsm transition --event finish_execute`. -The exact event names are read from `iq fsm state` at runtime (Principle I — the -skill says "use the event `iq fsm state` lists", never hardcodes a guess). - -## R3 — Skill format, naming, discovery - -- **Decision**: `iq-` (`iq-analyze`, `iq-plan`, `iq-execute`), one dir per - skill with `SKILL.md`, frontmatter `name` + `description` — identical to the - existing research/legion/kritik skills and to spec-kit's `speckit-*` skills. -- **Discovery**: deployed globally by `iq host get` into `~/.config/opencode/skills/` - and `~/.claude/skills/`; both hosts auto-discover (already proven). No new plumbing. - -## R4 — Brevity budget (SC-004) - -- **Decision**: the methodology core (Goal + Steps + Done-when) ≤ ~40 lines; - the artifact template is referenced/embedded separately and not counted against - the "reasoned-over" core. Mechanics are one-line `iq` commands. -- **Rationale**: spec-kit's effective core (`plan.md` outline) is ~7 steps; - brevity is what a weak model can actually follow. - -## R5 — Generation timing - -- **Decision**: generate at **deploy time** via `SkillBuilder` (mirrors how - `AgentBuilder` assembles the agent at deploy), reading the FSM state contract + - APE prompt + artifact template. Keeps generated skills out of source control and - always in sync with the installed CLI version. (A `--check` mode can fail CI if - a committed snapshot drifts, if desired later.) - -All NEEDS CLARIFICATION resolved. diff --git a/specs/001-iq-phase-skills/spec.md b/specs/001-iq-phase-skills/spec.md deleted file mode 100644 index 9f3aa0ea..00000000 --- a/specs/001-iq-phase-skills/spec.md +++ /dev/null @@ -1,94 +0,0 @@ -# Feature Specification: iq-* phase skills - -**Feature Branch**: `282-sdd-iq-phase-skills` - -**Created**: 2026-06-25 - -**Status**: Draft - -**Input**: User description: "Per-phase skills (/iq-analyze, /iq-plan, /iq-execute) so a weak model — or a human — can run the Inquiry methodology step-by-step when the inquiry scheduler agent cannot autonomously drive the CLI. Generated from the existing FSM + APE contracts so they stay in sync; brief but highly functional." - -## User Scenarios & Testing *(mandatory)* - -### User Story 1 - Run ANALYZE manually/assisted when the model can't drive (Priority: P1) - -A developer is working on an inquiry cycle with a weak local model (e.g. qwen3-coder:30b) that ignores or mis-drives the `inquiry` scheduler agent. Instead of relying on the agent, the developer invokes the `/iq-analyze` skill. The skill guides them (or the weak model) through the ANALYZE phase: what the phase must accomplish, the exact `iq` commands to run, the required shape of `diagnosis.md`, and how to pass the gate — until `complete_analysis` succeeds. - -**Why this priority**: ANALYZE is the first phase and the one where weak models fail most (engagement + producing a gate-compliant diagnosis). Delivering just this skill already lets an accessible model complete the hardest step, which is the project's core goal. - -**Independent Test**: Invoke `/iq-analyze` on a repo at the ANALYZE state and follow it; success = a `diagnosis.md` that passes `iq fsm transition --event complete_analysis` (exit 0), without using the scheduler agent. - -**Acceptance Scenarios**: - -1. **Given** a cycle in the ANALYZE state, **When** the user follows `/iq-analyze`, **Then** they produce a `diagnosis.md` that passes the `complete_analysis` gate. -2. **Given** the gate rejects the diagnosis, **When** the user re-reads the skill's "Done when" checklist and the gate error, **Then** the skill tells them exactly what to fix (e.g. a missing evidence handle) and they pass on retry. - ---- - -### User Story 2 - Run PLAN and EXECUTE the same way (Priority: P2) - -The developer continues the cycle with `/iq-plan` (produce a gate-passing `plan.md` with executable checks) and `/iq-execute` (implement the plan phase-by-phase under its constraints), each a self-contained guide for that phase. - -**Why this priority**: Completes the manual/assisted path end-to-end, but depends on the pattern proven by P1. - -**Independent Test**: From a repo in PLAN (resp. EXECUTE), follow `/iq-plan` (resp. `/iq-execute`); success = the phase's gate passes. - -**Acceptance Scenarios**: - -1. **Given** a cycle in PLAN, **When** the user follows `/iq-plan`, **Then** `plan.md` passes the PLAN→EXECUTE gate. -2. **Given** a cycle in EXECUTE, **When** the user follows `/iq-execute`, **Then** the plan's phases are implemented and the EXECUTE gate passes. - ---- - -### User Story 3 - Skills stay in sync with the methodology automatically (Priority: P2) - -A maintainer changes an FSM state contract or an APE operator prompt. The `iq-*` skills are regenerated from those sources, so they never drift from the actual gates and methodology the CLI enforces. - -**Why this priority**: Without this, the skills rot the moment the contracts change — the same drift problem the single-source firmware solved. It is what makes the skills trustworthy over time. - -**Independent Test**: Change a phase contract (e.g. add a required section to the diagnosis), regenerate, and confirm the corresponding `iq-*` skill reflects the change without manual edits. - -**Acceptance Scenarios**: - -1. **Given** a changed phase contract, **When** the skills are rebuilt, **Then** the affected `iq-*` skill content matches the new contract. - ---- - -### Edge Cases - -- The user invokes `/iq-analyze` when the cycle is **not** in ANALYZE → the skill must tell them to run `iq fsm state` and which phase/skill actually applies, rather than proceeding wrongly. -- No cleanroom/cycle exists yet → the skill points to the `iq`/`gh`/`git` commands that start a cycle (start/end are not their own skills). -- The gate keeps failing → the skill's "Done when" checklist + the gate's own error message must be enough to self-correct (no scheduler agent needed). - -## Requirements *(mandatory)* - -### Functional Requirements - -- **FR-001**: The system MUST provide three phase skills — `iq-analyze`, `iq-plan`, `iq-execute` — invocable on demand (no scheduler agent required). Start and end are handled by direct `iq`/`gh`/`git` commands, not dedicated skills. -- **FR-002**: Each skill MUST state the phase's **goal** and methodology (from the FSM state contract + the phase's APE operator), the exact **`iq` commands** for that phase (the mechanics), the required **artifact shape** (e.g. `diagnosis.md` sections + evidence-handle rule), and a **"Done when"** checklist tied to the phase's gate. -- **FR-003**: Each skill MUST be **brief but highly functional** — the methodology core (the part a human/weak model reasons over) MUST be short; mechanical work is delegated to `iq` commands and the artifact shape to a referenced template. -- **FR-004**: The skills MUST be **generated from the existing contracts** (FSM state contracts and APE operator prompts) at build time, so they cannot drift from the gates the CLI enforces. -- **FR-005**: The skills MUST be deployable to a host alongside the existing skills (research/legion/kritik) via the established deploy path. -- **FR-006**: A skill invoked in the wrong phase MUST direct the user to `iq fsm state` and the applicable phase, not produce a wrong artifact. - -### Key Entities *(include if feature involves data)* - -- **Phase skill (`iq-`)**: a self-contained, on-demand guide for one FSM phase; attributes — goal, methodology, `iq` commands, artifact shape, gate "Done when". -- **Phase contract**: the existing source of truth for a phase — the FSM state contract (mission, constraints, required artifacts, gate) + the phase's APE operator prompt (cognitive method). The skill is derived from it. - -## Success Criteria *(mandatory)* - -### Measurable Outcomes - -- **SC-001**: A user (or weak model) with no scheduler agent can take a cycle from ANALYZE to a passing `complete_analysis` gate using only `/iq-analyze` — measured by gate exit code 0. -- **SC-002**: With the three skills, a cycle can be driven manually/assisted through ANALYZE → PLAN → EXECUTE with every phase gate passing, without invoking the `inquiry` scheduler agent. -- **SC-003**: A weak local model assisted by `/iq-analyze` passes the ANALYZE gate at a **higher rate** than the same model driving the scheduler agent unaided (the manual skill is more followable than autonomous orchestration). -- **SC-004**: The methodology core of each skill stays short (a human can read and act on it in under ~2 minutes per phase), with mechanics offloaded to `iq` commands. -- **SC-005**: A change to a phase contract is reflected in the regenerated skill with zero manual skill edits (no drift). - -## Assumptions - -- The host (OpenCode, Claude Code) discovers on-demand skills from its skills directory, as the existing research/legion/kritik skills already prove. -- The existing `iq` CLI gates (e.g. `complete_analysis`) already verify artifacts and return actionable errors — the skills lean on them rather than re-implementing validation. -- Start/end of a cycle are adequately covered by direct `iq`/`gh`/`git` commands, so `iq-start`/`iq-cleanroom`/`iq-end` skills are out of scope for v1. -- The three current FSM phases map cleanly to three skills; additional states (IDLE, END, EVOLUTION) are out of scope for v1. diff --git a/specs/001-iq-phase-skills/tasks.md b/specs/001-iq-phase-skills/tasks.md deleted file mode 100644 index ebabc261..00000000 --- a/specs/001-iq-phase-skills/tasks.md +++ /dev/null @@ -1,74 +0,0 @@ -# Tasks: iq-* phase skills - -**Input**: Design documents from `/specs/001-iq-phase-skills/` -**Prerequisites**: plan.md, spec.md, research.md, data-model.md, contracts/skillbuilder.md -**Tests**: included (requested by spec — gate consistency + the SC-003 experiment). -**Organization**: by user story; US1 is the MVP. - -## Format: `[ID] [P?] [Story] Description` -- **[P]**: parallelizable (different files, no dependency). Paths are repo-relative. - ---- - -## Phase 1: Setup - -- [ ] T001 Create `code/cli/assets/artifacts/` directory for first-class artifact templates (research R1). - -## Phase 2: Foundational (blocking — no story can start until done) - -- [ ] T002 Create `code/cli/lib/hosts/skill_builder.dart`: `SkillBuilder(Assets)` with `build(String phase)` and `phaseSkillNames` per `contracts/skillbuilder.md`; reads `fsm/states/.yaml`, `apes/.yaml`, `artifacts/.template.md`. -- [ ] T003 [P] In `skill_builder.dart`, encode the phase→operator→artifact→gate map (data-model.md): analyze→socrates→diagnosis.md→complete_analysis; plan→descartes→plan.md→approve_plan/go_execute; execute→ada→(code)→finish_execute. -- [ ] T004 [P] Create `code/cli/test/skill_builder_test.dart` scaffold. - -**Checkpoint**: builder + test harness exist. - ---- - -## Phase 3 (US1, P1): `iq-analyze` end-to-end — MVP - -**Goal**: a user/weak model takes a cycle from ANALYZE to a passing `complete_analysis` using only `/iq-analyze`. -**Independent test**: deploy, follow `/iq-analyze`, gate exits 0 — no scheduler agent. - -- [ ] T005 [P] [US1] Author `code/cli/assets/artifacts/diagnosis.template.md` — Evidence/Hypotheses/Constraints/Open Questions with the "every Evidence bullet carries a re-checkable handle" rule, consistent with the `complete_analysis` gate. -- [ ] T006 [US1] Implement `SkillBuilder.build('analyze')`: assemble `SKILL.md` (frontmatter `iq-analyze`; `## Goal` from `analyze.yaml`; `## Steps` = `iq fsm state` → `iq ape prompt --name socrates` → write diagnosis → `iq fsm transition --event complete_analysis`, "use only events `iq fsm state` lists"; `## Artifact` = embedded template; `## Done when` checklist). Core ≤ ~40 lines (SC-004). -- [ ] T007 [US1] Wire `HostDeployer` to write `/iq-analyze/SKILL.md` (via SkillBuilder) during `iq host get`, beside the static skills. -- [ ] T008 [US1] Test (`skill_builder_test.dart`): `build('analyze')` contains `name: iq-analyze`, `iq fsm state`, the diagnosis template headers, `complete_analysis`, and a Done-when checklist; core line count ≤ budget. -- [ ] T009 [US1] Test: the `diagnosis.template.md`, written into a real cycle, **passes** `iq fsm transition --event complete_analysis` (template ⇄ gate consistency). -- [ ] T010 [US1] Test/verify: `iq host get` deploys `iq-analyze/SKILL.md` into the host skills dir (real-binary check). - -**Checkpoint**: `/iq-analyze` shippable on its own. - ---- - -## Phase 4 (US2, P2): `iq-plan` + `iq-execute` - -- [ ] T011 [P] [US2] Author `code/cli/assets/artifacts/plan.template.md` (executable-check rule, per-phase), consistent with the PLAN→EXECUTE gate. -- [ ] T012 [US2] Implement `SkillBuilder.build('plan')` and `build('execute')` (reusing the shared assembly from T006). -- [ ] T013 [US2] Deploy `iq-plan` + `iq-execute` (extend T007 to all `phaseSkillNames`). -- [ ] T014 [P] [US2] Tests: plan/execute skill content + `plan.template.md` ⇄ PLAN gate consistency. - -**Checkpoint**: full manual ANALYZE→PLAN→EXECUTE path. - ---- - -## Phase 5 (US3, P2): no-drift (generated from contracts) - -- [ ] T015 [US3] Test: editing a contract (e.g. add a required section to `analyze.yaml`) changes `build('analyze')` output with **no** manual skill edit (SC-005 / Principle V). - ---- - -## Phase 6: Polish & cross-cutting - -- [ ] T016 [P] **SC-003 experiment**: conducted run — same task on qwen-30b, once driving the scheduler agent unaided, once following `/iq-analyze`; measure `complete_analysis` pass rate; record evidence (per Constitution: evidence, not inference). -- [ ] T017 [P] CHANGELOG entry + version bump (`pubspec.yaml`/`version.dart`/site badge) + `quickstart.md` link from docs. -- [ ] T018 `dart analyze` clean + full `dart test` green + build & deploy verify with the real binary; open PR. - ---- - -## Dependencies & order -- Phase 2 (T002–T004) blocks all stories. -- **US1 (P1)** is the MVP and must land first; **US2** reuses US1's assembly; **US3** is a test over the builder; Phase 6 polish last. -- Parallel within a story: `[P]` tasks touch different files (e.g. T005 template vs T008/T009 tests). - -## Implementation strategy -Ship **US1 (`/iq-analyze`) as the first PR** (MVP — proves the 3-layer pattern + template⇄gate consistency + the SC-003 experiment on the hardest phase). Then US2 (plan/execute) and US3 (no-drift) in a follow-up. diff --git a/specs/002-iq-specification/spec.md b/specs/002-iq-specification/spec.md deleted file mode 100644 index e5001b7e..00000000 --- a/specs/002-iq-specification/spec.md +++ /dev/null @@ -1,105 +0,0 @@ -# Feature Specification: iq-specification (the QA specification phase) - -**Status**: Draft · **Branch**: `spec/iq-specification` (planned) · **Date**: 2026-06-27 - -## Summary - -`iq-specification` is the manual, QA-facing **specification phase** of the Inquiry -method. It turns a raw requirement (arriving as email, documents, chat, a -message) into a **healthy, coherent, actionable, well-granulated specification** -plus the GitHub issues that decompose it — deciding by **evidence from -throwaway experiments, not by inference**. It is **independent** of the -developer cycle (analyze → plan → execute): it *precedes* it and produces the -issues each developer cycle then consumes. - -It mirrors the user's engineering handbook (`cacsi-dev/handbook` → -`especificacion.md`), which now references this Inquiry methodology instead of a -separate CLI. - -## Team / workspace separation - -| Role | Skill(s) | Workspace | Level | -|---|---|---|---| -| **QA** | `iq-specification` | `requisitions//` | WHAT / why (requirement) | -| **Dev** | `iq-analyze` → `iq-plan` → `iq-execute` | `cleanrooms//` + code | HOW (implementation) | - -Handoff = the GitHub issues. (Mirrors spec-kit's "spec = what/why, never how".) - -## User Scenarios - -### User Story 1 — Turn a raw requirement into a specification (Priority: P1) -A QA analyst receives a requirement by email with attachments. They run -`iq-specification`, which scaffolds `specification.md`; they investigate (running -throwaway experiments where a decision needs evidence) and fill it with user -stories, Given-When-Then acceptance criteria, a testing strategy, explicit -scope, and a Decisions(evidence) section. -**Acceptance**: a filled `specification.md` that passes the `specification_ready` -quality gate. - -### User Story 2 — Decompose into well-scoped issues (Priority: P1) -From the specification, the analyst derives one `issue-.md` per issue -(tracked in-repo as the platform-independent source of truth), and the skill -prints the `gh` commands to create them — **after** a duplicate check -(`gh issue list --search`); if a too-similar issue exists, it edits instead. -**Acceptance**: one `issue-.md` per identified issue + ready `gh` commands; -no duplicate issues created. - -### User Story 3 — Decide by evidence, not inference (Priority: P2) -Where a scope/AC decision is uncertain, the analyst runs a throwaway probe (a DB -query, a container run, an API call) and records the decision with its -re-checkable evidence handle in the Decisions(evidence) section. -**Acceptance**: each key decision in the spec cites a re-checkable handle. - -## Requirements (functional) - -- **FR-1**: A new `iq-specification` skill, assembled from the contracts, using - the **DEWEY** operator. Steps: gather the raw requirement from all sources → - experiment for evidence → fill `specification.md` → derive `issue-.md` - per issue → duplicate-check + print `gh` commands → present to human and stop. -- **FR-2**: The CLI scaffolds `requisitions//specification.md` from a - single-source template (the hands), in the selected language. -- **FR-3**: **Bilingual** templates — `en` (default) and `es` via a `--lang` - flag. (Scope: the QA spec only; dev-cycle artifacts stay English because the - gates parse English.) -- **FR-4**: The template mirrors the handbook `especificacion.md`: Metadata, - User Stories (As a / I want / So that) + Acceptance Criteria (Given-When-Then), - Testing Strategy (Unit/Integration/E2E), Explicit Scope, **Decisions - (evidence)**, Annexes. -- **FR-5**: A `specification_ready` gate: each user story has ≥1 Given-When-Then - AC, explicit scope is present, a testing strategy is present, ≥1 issue is - derived, and each Decisions bullet carries a re-checkable handle. -- **FR-6**: Issues live as `issue-.md` in the repo (source of truth); - GitHub issues are a projection synced via the printed `gh` commands. -- **FR-7**: Methodology: `requisitions//` lands on `main` via a doc-PR - from `spec/` (the QA review gate); main is branch-protected. - -## Success Criteria - -- **SC-1**: A QA analyst (or capable model) using `iq-specification` produces a - `specification.md` that passes `specification_ready`, plus the issue files + - `gh` commands, without touching the developer cycle. -- **SC-2**: The Inquiry template and the handbook `especificacion.md` are - structurally identical (the handbook references Inquiry). -- **SC-3**: Every key decision in a produced spec carries a re-checkable - evidence handle (evidence over inference, measured on a sample). - -## Open design decisions (resolve before wiring) - -1. **Skill shape**: `iq-specification`'s step structure differs from the - analyze/plan/execute 5-step shape (it has experiment + issue-derivation + - gh-command steps). Extend `SkillBuilder` with a third shape, or give - `iq-specification` its own builder path. -2. **Goal source**: there is no `fsm/states/specification.yaml` (it is - standalone). Add a small `specification` contract (goal + method) vs reuse - `idle.yaml`/`dewey`. -3. **`--lang` plumbing**: a CLI flag vs a per-project config setting; default - `en`. -4. **Gate home**: `specification_ready` as a standalone check (no FSM state) vs - a mini-FSM for the specification phase. - -## Out of scope - -- Auto-creating GitHub issues (manual mode prints `gh` commands; the FSM/agent - mode could auto-create later). -- Bilingual dev-cycle artifacts (`diagnosis.md`, `plan.md` stay English). -- An `iq start` command (separate work).