Skip to content

Sessions lost on ungraceful exit — no incremental .jsonl flush #1

Description

@todie

Summary

omp allocates a session ID eagerly (before the agent loop runs) but does not incrementally persist the session .jsonl — it appears to commit only on clean exit. When a pane dies ungracefully (OOM-kill, SIGKILL, parent death, WSL2 VM reboot), the entire session transcript is lost despite having an allocated session ID.

Evidence

Surfaced during an omp-context audit on 2026-07-17. A fresh pane was spawned via omp @/tmp/darius_todo_seed.md to work on OPS-647. The pane worked for 16 minutes, then died. Zero bytes of its session were persisted.

Timeline (from ~/.omp/logs/omp.2026-07-17.log, times local America/New_York):

Time Event
23:33:36 Pane spawned (pid 2734200), TSSR rules registered, yaml-frontmatter warnings for rust-skills plugin (non-fatal)
23:35:11 title-generator: success — session ID 019f7349-5e6d-7000-92e7-0e6ed4cefb42 allocated. Title: "Authorize SSH key in darius-vm microVM"
23:37:51 First mid-run todo nudge fired (7 incomplete todos — pane was actively working the task)
23:41:49 Second todo nudge
23:47–23:49 Repeated ui.loop-blocked warnings on pids 59920, 30275, 2734200 (UI-loop stalls 250–1824 ms across all 3 live panes)
23:49:44 Last log line for pid 2734200 (ui.loop-blocked)
after silence — no crash signature, no exit, no SIGTERM trace
now kill -0 2734200 fails; zero .jsonl files modified in last 20 min across all of `~/.omp/`

No .jsonl exists on disk for session 019f7349-5e6d-7000-92e7-0e6ed4cefb42. The session ID exists in the log but no file backs it.

Root cause (hypothesis)

omp writes session .jsonl only on clean exit. An ungraceful death — OOM under memory pressure (WSL2 + 2.5 GB kokoro TTS + two ~750 MB long-lived panes + transient fresh spawns) — loses everything post-title-generator.

This is especially severe on WSL2, which is prone to OOM-kill under LLM-aware workload spikes. A single fresh spawn under memory pressure can lose its entire transcript; long-lived panes make that pressure worse for any new pane.

Proposed fix

Flush session .jsonl incrementally rather than only on clean exit. Options ranked by safety vs. implementation cost:

  1. Append + fsync every N messages (or every tool call). Robust against SIGKILL/OOM. Highest I/O cost but simplest correctness argument.
  2. Periodic flush (timer) every T seconds. Lower I/O but can lose the tail up to T.
  3. Flush on every tool-call boundary (between tool input and the next assistant turn). Best point-in-time granularity where the work actually happens.

A (3)-style implementation is probably the right balance — tool-call boundaries are natural delimitation points and align with when the agent has produced observable state worth persisting.

Repro

Hard to reproduce deterministically (requires an OOM-kill or SIGKILL mid-session). Easier synthetic repro:

  1. Start an omp pane on a moderately long task (seed file works).
  2. Wait until title-generator: fires and confirms a session ID in the omp log.
  3. kill -9 <pane pid> before the task completes.
  4. Observe no .jsonl exists for that session ID in ~/.omp/agent/sessions/<namespace>/.

Environment

  • omp v17.0.1 (@oh-my-pi/pi-coding-agent)
  • WSL2 (Linux 6.6.87.2-microsoft-standard-WSL2), 8-core Ryzen 9800X3D
  • Config: yolo approval, hashline edits, snapcompact (24k keep / 20k reserve, mid-turn)
  • MCP: only engram wired
  • Spawned via omp @<seedfile>.md
  • Model lane: llm/neuralwatt/glm-5.2:max (default) via LiteLLM gateway

Discoverer

ctodie — surfaced during the 2026-07-17 omp-context audit. Audit notes available on request.

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions