Skip to content

Run full-domain LongBench-v2 CodeQA (or a larger sample) to test the scale where direct baselines actually fail #172

Description

@shanev

Context

The published LongBench-v2 CodeQA result (benchmarks/results/longbench-v2-codeqa-20-2026-07-17/, PR #171) is a disclosed, cost-bounded 20-of-50 stratified subsample of the Code Repository Understanding domain, not the full 50-task domain — full-domain cost was estimated at $460-610 (one pilot task alone cost $3.30 on 659K tokens), disproportionate relative to the other #166 benchmark families ($20-30 each).

At n=20, droste's own README already discloses the direct-vs-droste comparison is underpowered: "the two-task score difference from direct-sol is the observed result, not evidence of a population-level separation." droste-terra-luna ties direct-terra and trails direct-sol by 10 points at this sample size — a real, current result, but not a statistically meaningful one given how few tasks separate the arms.

What's needed

Run LongBench-v2 CodeQA at a larger sample size (or the full 50-task domain) once funding allows, to get a droste-vs-direct-model comparison with enough statistical power to actually distinguish the arms, rather than the current small, cost-bounded sample.

Estimate

~$460-610 for the full 50-task x 3-arm domain (established during the original pilot). A partial expansion (e.g. 30-35 tasks, weighted toward the "long" stratum) could meaningfully increase statistical power at a fraction of that cost — worth estimating precisely before committing to full-domain.

Not blocking

This does not block droste#166's current publication (now fully complete) — see PR #171 for the existing, honestly-disclosed 20-task result.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions