A personal monorepo for agent skills and supporting scripts.
Install the plugin when you need any composed workflow. It packages all nine
skills together so stable-name dependencies such as implement-ticket,
review-code-change, and babysit-pr are available in the same fresh session.
Add the repository marketplace, install the plugin, and start a fresh session:
/plugin marketplace add shaug/agent-scripts
/plugin install agent-scripts@agent-scripts
The equivalent non-interactive commands are:
claude plugin marketplace add shaug/agent-scripts
claude plugin install agent-scripts@agent-scriptsCodex CLI 0.117 or newer can add the repository marketplace and install the same plugin:
codex plugin marketplace add shaug/agent-scripts
codex plugin add agent-scripts@agent-scriptsThe repository marketplace also appears in the Codex plugin browser. Start a fresh Codex task after installation so the complete skill set is loaded.
Standalone-capable skills can still be installed independently:
npx skills add shaug/agent-scriptsCodex users can also invoke $skill-installer with this repository. Prefer the
plugin for implement-epic, implement-ticket, babysit-pr,
review-code-change, or carve-changesets: installing one of those entrypoints
alone intentionally leaves required stable-name dependencies unavailable and
causes the workflow to fail closed.
skills/— skill folders, each containing aSKILL.mdand bundled resourcesreview-suite/— canonical code-review contracts, validators, raw evaluation fixtures, and the result-blind replay evaluator shared by repository-owned review skillsjustfile— common tasks for testing, validation, and formatting
Current reusable agent skills:
skills/babysit-pr— monitor one existing GitHub pull request through current-head CI, feedback, repository-owned re-review, mergeability, and an explicitly authorized completion policyskills/implement-ticket— implement exactly one standalone ticket or named epic child through isolated execution and initial repository-owned review, publish one ordinary PR throughbabysit-pror an explicitly authorized carved stack throughcarve-changesets, then verify tracker, mainline, and cleanup outcomes; this is the canonical owner of generic single-ticket execution rules consumed byimplement-epicskills/implement-epic— traverse live GitHub or Linear epic graphs and delegate each selected child toimplement-ticket, then refresh graph state and verify separately authorized epic closeoutskills/carve-changesets— recompose a review-ready source branch into a stateless chain, review each changeset throughreview-code-change, and delegate each published PR lifecycle tobabysit-prskills/review-code-change— orchestrate the repository-owned review lenses into one evidence-bound, deduplicated verdictskills/review-correctness— find material behavioral, security, compatibility, data-integrity, and validation failures in a code changeskills/review-solution-simplicity— challenge whole-solution machinery that is not justified by real requirements or repository constraintsskills/review-code-simplicity— reduce local cognitive load through behavior-preserving reuse, DRY, control-flow, and test simplificationskills/review-fix-loop— take cooperative ownership of an existing committed candidate, run the complete repository review suite in a fresh read-only subagent by default, apply ticket-scoped fixes in isolated attempt worktrees, and repeat until review converges or a bounded stop condition is reached; supports alocal_commitpolicy (every fix stays local — the operator publishes them separately) and anupdate_prpolicy (one expected-old, fast-forward-only Git push immediately after convergence). Standalone today: no existing caller skill invokes it yet — see issues #103, #104, and #105 for the trackedimplement-ticket/babysit-pr/carve-changesetsmigration follow-ups.
The composed implementation dependency chain is:
implement-epic
└── implement-ticket
├── review-code-change # initial candidate review
├── babysit-pr # ordinary single-PR lifecycle
│ └── review-code-change # after a head-changing fix
└── carve-changesets # authority-gated oversized path
├── review-code-change # each exact changeset
└── babysit-pr # each changeset PR lifecycle
└── review-code-change # after a head-changing fix
carve-changesets
├── review-code-change # direct per-changeset review
└── babysit-pr # each published PR lifecycle
└── review-code-change # after a head-changing fix
Compatible runtimes may provide named subagents or equivalent isolated
implementation and review contexts. Files under each skill's agents/ directory
(openai.yaml for OpenAI runtimes, claude-code.md for Claude Code) are
optional discovery and adapter metadata, not part of the skills' portable
contracts.
Each review skill bundles a verbatim copy of the canonical review-suite
contract and schemas under its own references/review-suite/ directory so the
skill remains self-contained when installed elsewhere. Edit only the canonical
files and refresh the copies with:
just sync-contractsRun the core checks:
just checkRun skill-specific tests:
just test-carve-changesets
just test-review-suite
just test-babysit-pr
just test-implement-ticket
just test-implement-epic
just test-review-fix-loop
just eval-implement-ticket
just eval-implement-epic
just eval-review-fix-loopValidate a review packet and result together:
python3 review-suite/scripts/validate.py pair packet.json result.jsonRun deterministic local evaluation harnesses without an agent runtime:
just eval-carve-changesets
just eval-implement-ticket
just eval-implement-epic
just eval-review-fix-loopThe carve-changesets command first runs its objective integration self-test, which checks clean-tree, plan, immutable-source, chain, equivalence, and validation invariants. It then runs peer-shaped judgment cases through a fresh process for each result-blind packet. The bundled executor is a deterministic simulation of a compliant runtime, not a model evaluation. Pass any compatible stdin/stdout JSON adapter with:
just eval-carve-changesets-executor "python3 /path/to/adapter.py"The ticket-composition evaluator starts a fresh process for each case, with
fixture identity and grader expectations withheld. Its shared corpus targets
both implement-ticket and implement-epic; just eval-implement-epic filters
the same validated corpus to epic packets. Case artifacts carry raw, live-shaped
workflow state, including criterion-specific acceptance evidence bound to the
candidate or deployed SHA. The harness grades obligation mapping and
terminal-state selection. Its bundled reference executor is a deterministic
simulation of a compliant runtime, not a model. To forward-evaluate a real agent
runtime, pass its stdin/stdout JSON adapter through
scripts/evals/run_forward.py --executor and retain captured observations with
--output-dir. A Claude Code headless adapter is bundled:
just eval-implement-ticket-claudejust eval-review-fix-loop drives review-fix-loop's own cross-cutting,
result-blind corpus: twenty scenarios that run the real local_commit/
update_pr engine against disposable Git repositories with scripted
reviewer/decide/apply_fix fixtures (no subprocess boundary, no model call, no
network) and grade the result against independently derived Git evidence. See
its README for the corpus contract.
review-suite/evals/ is the canonical evaluator that measures what the review
skills actually do when a real agent runtime executes them repeatedly, rather
than asserting quality through expected JSON. It changes no review behaviour.
See its README for the protocol, the corpus
contract, and the grading interface.
Three commands cover it:
just test-review-suite # deterministic tests, no runtime
just audit-review-corpus # corpus integrity, no runtime
just eval-review-suite '<executor command>' # the only one that may cost moneyjust test includes the deterministic evaluator tests and never launches a paid
runtime. just eval-review-suite is deliberately absent from test, lint,
and check. The bundled real-runtime adapter needs the claude CLI on PATH:
just eval-review-suite "python3 review-suite/scripts/evals/claude_executor.py"Each attempt starts a fresh process and receives one result-blind JSON request:
the target skill text, the raw review packet, the review contracts, and public
run identity. Expected findings, private grader labels, provenance, and even the
case name stay out of the payload, and just audit-review-corpus proves it by
inspecting the complete payload every case would produce. Spawn, timeout,
runtime, oversized-output, malformed, and protocol failures are reported
separately from a valid review and are never scored as clean.
The bundled fixture_executor.py is a deterministic simulation of a compliant
reviewer, not a model. Its runs are marked simulation, so no baseline report
can be produced from them. Run the runner directly for repeated attempts,
per-attempt records, and the aggregate report:
python3 review-suite/scripts/evals/runner.py \
--executor "python3 review-suite/scripts/evals/claude_executor.py" \
--runs 5 --report-out out/report.jsonCases are grouped into strata under review-suite/evals/strata/. A stratum
is the unit of valid comparison — same target skill, same declared dependency
closure, same runtime and model, same kind of ground truth — and each directory
is a complete corpus declaring its own target. just audit-review-corpus
discovers every one of them.
The frozen v1 baseline record lives in review-suite/evals/baseline/v1/: the
immutable configuration, the unscored pilot's per-stratum cost and latency
envelope, a per-stratum cost-ceiling proposal built from those numbers, the
grader calibration and adjudication record, a plan for satisfying the
two-independent-adjudication gate, the ground-truth sourcing and sanitization
record, and the baseline limitations.
Two things are deliberately still outstanding, and the configuration record says so rather than implying otherwise: the scored strata are declared but not populated, and no scored run may launch until the repository owner preregisters a per-stratum cost ceiling and each private expectation has two independent adjudications. Read the limitations record before quoting any figure — in particular, the connector stratum is deferred, not satisfied, so connector-escape recall has never been measured and no human-review figure may be reported as a connector figure.
The v2 scored closeout lives in review-suite/evals/v2/: the preregistered gate
manifest and decision record from #59, and #57's own scored ablation
comparison — three independently-measured s1-correctness-orchestrator
configurations (each required pass isolated, then both together) plus the
reused, non-regressed s2/s3 figures, each checked against the settled
per-case gate. It reports what clean proves, what it does not prove (in
particular, whether a passing numeric gate demonstrates a mechanism's own unique
causal contribution — it does not, on the evidence recorded there), and a
recommendation, not a removal decision, per mechanism. See
its README for the full document set.
Following that evidence, the repository owner removed the verification-
sufficiency pass (#93): neither #57's ablation matrix nor #89's harder-
case validation (discriminating-case-validation.md) found demonstrated value
for it, and it carried a confirmed, twice-reproduced false-positive regression
on session-continuation-summary when run without the traversal pass. The
traversal (consumer/impact) pass and consumer_impact_evidence were unaffected
— that pass showed a real, reproducible gap in the same validation. The shared
review-result contract advanced 1.3 → 1.4 to drop
verification_sufficiency_evidence; see
review-suite/evals/v2/VERIFICATION-SUFFICIENCY-REMOVAL.md
for the full decision record.
review-suite/evals/curation/ turns a newly adjudicated connector finding into
a versioned curation record, and turns a group of curation records into a
conservative, evidence-backed promotion decision: a corpus case only, a global
rubric change, a repository-instruction change, or nothing. It is infrastructure
only, proven with synthetic fixtures, and never touches baseline/v1/ or v2/.
See its README for the disposition
vocabulary, the mechanical disclosure guardrail that fails closed on a leaked
source identifier, and the promotion rules.
just audit-review-curation # curation and promotion-decision integrity, no runtime- Python 3.11+
- Git
- GitHub CLI (
gh) 2.37 or newer for PR workflows (the babysit-pr watcher requiresgh pr checks --json) skills-refonPATHfor skill validation (optional but recommended)
Each skill should be runnable and testable in isolation. Prefer adding tests
under the skill’s own scripts/tests/ directory and wire them into the
justfile.