Skip to content

[2/6][rollout] feat: container contract, workflow output types, in-container runner - #284

Merged
mathewjhan merged 1 commit into
mainfrom
mathew/container-contract
Aug 6, 2026
Merged

[2/6][rollout] feat: container contract, workflow output types, in-container runner#284
mathewjhan merged 1 commit into
mainfrom
mathew/container-contract

Conversation

@mathewjhan

@mathewjhan mathewjhan commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Part of splitting #272 into reviewable pieces (2/6).

What this adds

The data types and the in-container code that let a user's agent workflow run inside a task container and report results back to the host.

  • ContainerInput / ContainerResult (container/files.py): the two JSON files exchanged with the container. The host stages container_input.json (rollout id, prompt, metadata, chat-completions URL, API key) before the agent starts; the agent phase writes container_result.json (status, error, workflow output) when it ends. write_reward writes the reward file at the path Harbor's verifier reads (/logs/verifier/reward.json).
  • AgentWorkflowOutput (types/output.py): what a workflow's run() may return — one or more named message histories plus metrics. Returning nothing is also valid; the runner then collects the conversation recorded at the chat-completions endpoint.
  • RolloutStatus (types/sample.py): one status vocabulary used everywhere — queued, running, grading, success, failure, cancelled, unknown. RolloutSample gains extra_fields for structured diagnostics.
  • The runner (container/runner.py): the entrypoints that the generated bundle scripts call inside the container. agent_main reads the input file, runs the workflow, writes the result file, and saves the conversation as an ATIF trajectory.json next to the agent logs — the same location and format native Harbor agents use, so tools that read trajectories work on both. grader_main loads the agent's messages (from the result file, or from a trajectory file when the agent was not SDK code), runs the user's grader, and writes the reward. The grader reads its input from tests/ first: that copy may carry the ground-truth label, while the agent-phase copy has the label removed so the model can never read the answer.
  • messages_from_trajectory (container/trajectories.py): converts a trajectory document (ATIF steps, or a raw messages list from a native harness) into plain chat messages.

Example

A workflow and grader written against these types:

class MyWorkflow(AgentWorkflow):
    async def run(self, ctx: AgentWorkflowContext) -> None:
        agent = StrandsAgent(model=ctx.config.model, messages=ctx.prompt)
        await agent.invoke_async()
        # returning None is fine: the recorded conversation becomes the sample

class MyGrader(Grader):
    async def grade(self, ctx: GraderContext) -> None:
        answer = ctx.sample.messages[-1]["content"]
        ctx.set_reward(1.0 if answer.strip() == ctx.label else 0.0)

Inside the container, the bundle's console scripts call agent_main(MyWorkflow, config) and grader_main(MyGrader, config); everything else in this PR is the plumbing those two calls rely on.

@mathewjhan mathewjhan changed the title [rollout] feat: container contract, workflow output types, in-container runner [2/6][rollout] feat: container contract, workflow output types, in-container runner Aug 6, 2026
Base automatically changed from mathew/rollout-ttl-cache to main August 6, 2026 22:29
…er runner

Co-authored-by: Cursor <cursoragent@cursor.com>
@mathewjhan
mathewjhan force-pushed the mathew/container-contract branch from f63bfc8 to cb62e58 Compare August 6, 2026 22:29
@mathewjhan
mathewjhan merged commit 24b72ea into main Aug 6, 2026
1 check passed
@mathewjhan
mathewjhan deleted the mathew/container-contract branch August 6, 2026 22:31
mathewjhan added a commit that referenced this pull request Aug 6, 2026
…dles (#285)

Part of splitting #272 into reviewable pieces (3/6).

## What this adds

`osmosis_ai/packaging.py`: builds a standard Python wheel from a user's
rollout project so the project can be installed inside a task container
with one `pip install`.

- `build_bundle(project_dir, workflow=..., grader=...)` produces a wheel
containing the project's package plus a generated `bundle_main.py` shim.
The shim imports the user's classes directly (`from my_harness.solver
import MyWorkflow`) and exposes two console scripts, `<package>-agent`
and `<package>-grade`, which call the runner entrypoints from #284.
Nothing is resolved dynamically at runtime; the class binding happens at
build time.
- Wheels are cached by a content hash of the project files under the
user cache directory (`platformdirs`), so rebuilding an unchanged
project is free.
- `inspect_bundle(wheel)` reads the wheel's metadata with
`importlib.metadata` and returns the declared dependencies (keeping
environment markers, dropping extras-gated entries) plus the two script
names. The Harbor backend (next in the stack) uses this list to
pre-install dependencies into the task image.

Also includes the `bench_harness` fixture project the packaging tests
build against, and the `platformdirs` dependency.

## Example

```python
from osmosis_ai.packaging import build_bundle, inspect_bundle

wheel = build_bundle(
    Path("my_rollout_project"),
    workflow="my_harness.solver:MyWorkflow",
    grader="my_harness.grade:MyGrader",
)
info = inspect_bundle(wheel)
info.agent_script    # "my-harness-agent"
info.requirements    # ["strands-agents>=1.0", "httpx>=0.27", ...]
```

Inside a container, `pip install my_harness-0.1.0-py3-none-any.whl`
followed by running `my-harness-agent` executes the user's workflow with
no other setup.

Co-authored-by: Cursor <cursoragent@cursor.com>
mathewjhan added a commit that referenced this pull request Aug 6, 2026
Part of splitting #272 into reviewable pieces (4/6). Builds on #284
(container contract) and #285 (packaging).

## What this adds

`HarborBackendV2`: runs each rollout as a Harbor trial. The agent can be
either a user workflow (packaged into a wheel and installed in the
container at trial start) or a registered native Harbor agent
(`terminus-2`, `mini-swe-agent`, `oracle`) with the rollout endpoint
injected into its environment.

How a rollout flows through it:

1. **Task selection** (`tasks.py`): template mode uses one task
directory for every rollout; dataset mode routes by
`metadata["harbor_task_id"]` to a folder under `tasks_dir` (path escapes
rejected); `metadata["harbor_task"]` fetches a task from a local path,
git checkout, or registry package, with per-ref locks so concurrent
rollouts download once.
2. **Materialization**: the task is copied into a per-rollout directory;
the rollout's input file is staged; if the task has no `tests/` and a
grader exists, a `test.sh` is generated that installs and runs the
grader. The ground-truth label is staged only into `tests/`, which
Harbor uploads at verification time — the agent phase cannot read it.
3. **Image preparation**: `patch_dockerfile_with_sdk` appends a block to
the task's Dockerfile that installs a static `uv` binary and creates
`/opt/osmosis/venv` with the bundle's dependencies pre-installed.
Per-trial installs then only add the user's own code (`--no-deps`),
which cuts container startup from minutes to seconds. The patch is
deterministic, so identical tasks keep identical image content hashes
and share builds.
4. **Execution** (`harness_agent.py`): the installed agent uploads the
wheel, installs it into the venv, backfills an empty prompt from the
task's `instruction.md`, runs the agent script, and returns the result
through the trial's agent metadata.
5. **Callbacks**: the workflow-complete callback fires when verification
starts (agent phase over); the grader-complete callback fires at trial
end with the reward parsed from Harbor's verifier result. Callback
delivery failures are logged and never abort trial archival.
6. **Observability and lifecycle**: per-phase timings and failure phases
in every result (`diagnostics.py`), native-agent ATIF parsing with
secret redaction, artifact relocation, `prewarm()` /
`prewarm_lifespan()` to build task images before serving traffic,
`cancel_rollouts(ids | prefix | all)`, `rollout_status()` with terminal
outcomes retained in a `TtlCache` (#283), and admission control via
`max_queue_depth`.

## Example

```python
backend = HarborBackendV2(
    orchestrator=TrialQueue(n_concurrent=100),
    tasks_dir=Path("tasks"),           # 300 task folders
    task_mode="dataset",
    agent=MyWorkflow,                  # or agent="mini-swe-agent"
    workflow_config=my_config,
    environment_config=EnvironmentConfig(type=EnvironmentType.SKYPILOT),
)
app = create_rollout_server(
    backend=backend,
    lifespan=backend.prewarm_lifespan(task_ids=["task-0000"]),
)
```

A trainer then POSTs rollouts with `metadata={"harbor_task_id":
"task-0042"}`; each one runs in its own sandbox and reports back through
the callbacks.

Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants