Applied Research Engineer at Abundant (YC)
Research on agent evals — benchmarks, harnesses, and environments for frontier coding agents.
I design agent benchmarks and the sandboxed environments that grade them deterministically — clean reinforcement-learning signal and rigorous evaluation for frontier coding agents. My current focus is long-horizon autonomy: multi-hour library reproductions, full-stack product clones, and ML builds.
Co-author of SWE-Marathon, a benchmark of 20 multi-hour software-engineering tasks (code · paper).
SWE-Marathon is cited in the official model cards of:
agent evaluation RL environments long-horizon tasks execution sandboxes model reliability



