Boilerplate-free, reproducible machine-learning experiment sweeps —
a framework-agnostic sweep engine built on
hydra-zen,
with first-class PyTorch Lightning integration.
Decorate an experiment, sweep over parameters, and get the results back as a
labeled xarray.Dataset — not rows in a dashboard you have to export.
- Boilerplate-free — one decorator, no subclassing or callbacks; results come back as a labeled dataset you can slice and plot.
- Reproducible — every run captures its config + provenance, sweeps resume durably after a hard kill or cluster preemption, and auto-tuning pins a hardware-independent effective batch size.
- Scalable — the same task runs in-process, across cores, or on a multi-node SLURM cluster (validated on real GPU hardware); change only the launcher.
- Framework-agnostic — your task just returns a
dict, so scikit-learn, XGBoost, or any Python model works, not only PyTorch.
import mushin
@mushin.sweep
def experiment(lr, seed):
# ... train a model with this lr/seed, then evaluate it ...
acc = ... # your validation accuracy
return dict(accuracy=acc) # whatever you return becomes a data variable
ds = experiment.run(
lr=mushin.multirun([0.01, 0.1, 1.0]),
seed=mushin.multirun([0, 1, 2]),
) # 9 runs, returned as a labeled xarray.Dataset
# <xarray.Dataset> Dimensions: (lr: 3, seed: 3)
# Data variables: accuracy (lr, seed)
ds["accuracy"].mean("seed") # average over seeds, per learning rate
# Prefer pandas? One call gives a tidy table — no xarray required:
experiment.workflow.to_dataframe() # lr seed accuracy
# 0 0.01 0 ...The decorator version above runs as
examples/parallel_sweep.py;
examples/sweep_to_dataset.py is the same flow
written with the class API. Need the full tool (.failures, provenance, custom
analysis)? Drop to experiment.workflow, or subclass
MultiRunMetricsWorkflow.
Each links to a guide with a runnable example:
- Compare methods with statistics —
benchmark.compareruns a metric battery (torchmetrics) across seeds and returns a labeled dataset plus significance (scipy);Studyruns the training sweep and feeds it straight in. - Compare LLM systems — the same significance spine for LLM evals, with Holm-corrected p-values and an optional output cache.
- Built-in batteries — classification, segmentation, detection, regression, retrieval, image-quality, audio.
- Resilient & resumable sweeps —
on_error="nan", durableresume=Trueacross hard kills/preemption, and per-run provenance. - Parallel & out-of-process launchers —
run(..., launcher="joblib")or submitit; stdlib-picklable dispatch. - Multi-node & sharded training —
HydraDDP/HydraFSDPandpin_gpu_round_robinGPU packing, validated on a real SLURM cluster. - Auto-tuning —
tune_batch_size/tune_learning_rate, pinned for reproducibility. - Analyze from Claude Code (MCP) — an optional read-only MCP server to load and inspect completed runs.
pip install mushin-py # the sweep -> dataset core
pip install "mushin-py[eval]" # + compare, metric batteries, LLM eval, StudyThe PyPI distribution is mushin-py, but you import mushin (like
scikit-learn → sklearn). The eval extra adds the evaluation layer
(compare, the metric batteries, LLM evaluation, Study) and its heavier
dependencies (torchmetrics, scipy) — keeping the core install lean; accessing
those features without it raises a clear install hint. Other optional extras:
viz, netcdf, detection, image, audio, mcp (the battery extras imply
eval) — e.g. pip install "mushin-py[eval,viz]". Supported Python: 3.10 – 3.13.
mushin follows SemVer with pre-1.0 semantics: a minor bump (0.8 → 0.9) may
contain breaking changes, always listed with migration notes in the
changelog; patch releases never break. The public API is the
top-level mushin namespace plus the documented mushin.benchmark /
mushin.llm symbols — underscore modules are internal. The project's scope is
deliberately narrow: boilerplate-free, reproducible sweep → dataset, with
an optional evaluation layer behind the eval extra. It complements — and
does not replace — experiment trackers like W&B or TensorBoard (see the
workflows guide for
using both together).
If you use mushin in your research, please cite it (the concept DOI always resolves to the latest release):
@software{martinez_mushin,
author = {Martínez-Martínez, Josué},
title = {mushin: boilerplate-free, reproducible machine-learning experiment sweeps},
year = {2026},
version = {0.11.1},
doi = {10.5281/zenodo.21436444},
url = {https://github.com/martinez-hub/mushin}
}GitHub's "Cite this repository" button (from CITATION.cff)
generates this for you.
Issues and pull requests are welcome. Local development uses uv:
uv sync
uv run pytest tests/ --hypothesis-profile fast # tests (DDP test needs >=2 GPUs)
uv run ruff check . && uv run ruff format --check .Or make check (lint + format + spell + tests, as CI runs). See
CONTRIBUTING.md.
mushin is a maintained, standalone carve-out of the rai_toolbox.mushin
subpackage from MIT Lincoln Laboratory's
responsible-ai-toolbox
(no longer maintained). This is a fork/extraction, not a replacement endorsed by
MIT-LL; the original MIT copyright is retained (see LICENSE.txt).
The configuration engine it builds on, hydra-zen, is actively maintained by the
same group.