[2/6][rollout] feat: container contract, workflow output types, in-container runner - #284
Merged
Merged
Conversation
This was referenced Aug 6, 2026
…er runner Co-authored-by: Cursor <cursoragent@cursor.com>
mathewjhan
force-pushed
the
mathew/container-contract
branch
from
August 6, 2026 22:29
f63bfc8 to
cb62e58
Compare
JoyboyBrian
approved these changes
Aug 6, 2026
mathewjhan
added a commit
that referenced
this pull request
Aug 6, 2026
…dles (#285) Part of splitting #272 into reviewable pieces (3/6). ## What this adds `osmosis_ai/packaging.py`: builds a standard Python wheel from a user's rollout project so the project can be installed inside a task container with one `pip install`. - `build_bundle(project_dir, workflow=..., grader=...)` produces a wheel containing the project's package plus a generated `bundle_main.py` shim. The shim imports the user's classes directly (`from my_harness.solver import MyWorkflow`) and exposes two console scripts, `<package>-agent` and `<package>-grade`, which call the runner entrypoints from #284. Nothing is resolved dynamically at runtime; the class binding happens at build time. - Wheels are cached by a content hash of the project files under the user cache directory (`platformdirs`), so rebuilding an unchanged project is free. - `inspect_bundle(wheel)` reads the wheel's metadata with `importlib.metadata` and returns the declared dependencies (keeping environment markers, dropping extras-gated entries) plus the two script names. The Harbor backend (next in the stack) uses this list to pre-install dependencies into the task image. Also includes the `bench_harness` fixture project the packaging tests build against, and the `platformdirs` dependency. ## Example ```python from osmosis_ai.packaging import build_bundle, inspect_bundle wheel = build_bundle( Path("my_rollout_project"), workflow="my_harness.solver:MyWorkflow", grader="my_harness.grade:MyGrader", ) info = inspect_bundle(wheel) info.agent_script # "my-harness-agent" info.requirements # ["strands-agents>=1.0", "httpx>=0.27", ...] ``` Inside a container, `pip install my_harness-0.1.0-py3-none-any.whl` followed by running `my-harness-agent` executes the user's workflow with no other setup. Co-authored-by: Cursor <cursoragent@cursor.com>
mathewjhan
added a commit
that referenced
this pull request
Aug 6, 2026
Part of splitting #272 into reviewable pieces (4/6). Builds on #284 (container contract) and #285 (packaging). ## What this adds `HarborBackendV2`: runs each rollout as a Harbor trial. The agent can be either a user workflow (packaged into a wheel and installed in the container at trial start) or a registered native Harbor agent (`terminus-2`, `mini-swe-agent`, `oracle`) with the rollout endpoint injected into its environment. How a rollout flows through it: 1. **Task selection** (`tasks.py`): template mode uses one task directory for every rollout; dataset mode routes by `metadata["harbor_task_id"]` to a folder under `tasks_dir` (path escapes rejected); `metadata["harbor_task"]` fetches a task from a local path, git checkout, or registry package, with per-ref locks so concurrent rollouts download once. 2. **Materialization**: the task is copied into a per-rollout directory; the rollout's input file is staged; if the task has no `tests/` and a grader exists, a `test.sh` is generated that installs and runs the grader. The ground-truth label is staged only into `tests/`, which Harbor uploads at verification time — the agent phase cannot read it. 3. **Image preparation**: `patch_dockerfile_with_sdk` appends a block to the task's Dockerfile that installs a static `uv` binary and creates `/opt/osmosis/venv` with the bundle's dependencies pre-installed. Per-trial installs then only add the user's own code (`--no-deps`), which cuts container startup from minutes to seconds. The patch is deterministic, so identical tasks keep identical image content hashes and share builds. 4. **Execution** (`harness_agent.py`): the installed agent uploads the wheel, installs it into the venv, backfills an empty prompt from the task's `instruction.md`, runs the agent script, and returns the result through the trial's agent metadata. 5. **Callbacks**: the workflow-complete callback fires when verification starts (agent phase over); the grader-complete callback fires at trial end with the reward parsed from Harbor's verifier result. Callback delivery failures are logged and never abort trial archival. 6. **Observability and lifecycle**: per-phase timings and failure phases in every result (`diagnostics.py`), native-agent ATIF parsing with secret redaction, artifact relocation, `prewarm()` / `prewarm_lifespan()` to build task images before serving traffic, `cancel_rollouts(ids | prefix | all)`, `rollout_status()` with terminal outcomes retained in a `TtlCache` (#283), and admission control via `max_queue_depth`. ## Example ```python backend = HarborBackendV2( orchestrator=TrialQueue(n_concurrent=100), tasks_dir=Path("tasks"), # 300 task folders task_mode="dataset", agent=MyWorkflow, # or agent="mini-swe-agent" workflow_config=my_config, environment_config=EnvironmentConfig(type=EnvironmentType.SKYPILOT), ) app = create_rollout_server( backend=backend, lifespan=backend.prewarm_lifespan(task_ids=["task-0000"]), ) ``` A trainer then POSTs rollouts with `metadata={"harbor_task_id": "task-0042"}`; each one runs in its own sandbox and reports back through the callbacks. Co-authored-by: Cursor <cursoragent@cursor.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of splitting #272 into reviewable pieces (2/6).
What this adds
The data types and the in-container code that let a user's agent workflow run inside a task container and report results back to the host.
ContainerInput/ContainerResult(container/files.py): the two JSON files exchanged with the container. The host stagescontainer_input.json(rollout id, prompt, metadata, chat-completions URL, API key) before the agent starts; the agent phase writescontainer_result.json(status, error, workflow output) when it ends.write_rewardwrites the reward file at the path Harbor's verifier reads (/logs/verifier/reward.json).AgentWorkflowOutput(types/output.py): what a workflow'srun()may return — one or more named message histories plus metrics. Returning nothing is also valid; the runner then collects the conversation recorded at the chat-completions endpoint.RolloutStatus(types/sample.py): one status vocabulary used everywhere —queued,running,grading,success,failure,cancelled,unknown.RolloutSamplegainsextra_fieldsfor structured diagnostics.container/runner.py): the entrypoints that the generated bundle scripts call inside the container.agent_mainreads the input file, runs the workflow, writes the result file, and saves the conversation as an ATIFtrajectory.jsonnext to the agent logs — the same location and format native Harbor agents use, so tools that read trajectories work on both.grader_mainloads the agent's messages (from the result file, or from a trajectory file when the agent was not SDK code), runs the user's grader, and writes the reward. The grader reads its input fromtests/first: that copy may carry the ground-truth label, while the agent-phase copy has the label removed so the model can never read the answer.messages_from_trajectory(container/trajectories.py): converts a trajectory document (ATIF steps, or a raw messages list from a native harness) into plain chat messages.Example
A workflow and grader written against these types:
Inside the container, the bundle's console scripts call
agent_main(MyWorkflow, config)andgrader_main(MyGrader, config); everything else in this PR is the plumbing those two calls rely on.