Context
The published LongBench-v2 CodeQA result (benchmarks/results/longbench-v2-codeqa-20-2026-07-17/, PR #171) is a disclosed, cost-bounded 20-of-50 stratified subsample of the Code Repository Understanding domain, not the full 50-task domain — full-domain cost was estimated at $460-610 (one pilot task alone cost $3.30 on 659K tokens), disproportionate relative to the other #166 benchmark families ($20-30 each).
At n=20, droste's own README already discloses the direct-vs-droste comparison is underpowered: "the two-task score difference from direct-sol is the observed result, not evidence of a population-level separation." droste-terra-luna ties direct-terra and trails direct-sol by 10 points at this sample size — a real, current result, but not a statistically meaningful one given how few tasks separate the arms.
What's needed
Run LongBench-v2 CodeQA at a larger sample size (or the full 50-task domain) once funding allows, to get a droste-vs-direct-model comparison with enough statistical power to actually distinguish the arms, rather than the current small, cost-bounded sample.
Estimate
~$460-610 for the full 50-task x 3-arm domain (established during the original pilot). A partial expansion (e.g. 30-35 tasks, weighted toward the "long" stratum) could meaningfully increase statistical power at a fraction of that cost — worth estimating precisely before committing to full-domain.
Not blocking
This does not block droste#166's current publication (now fully complete) — see PR #171 for the existing, honestly-disclosed 20-task result.
Context
The published LongBench-v2 CodeQA result (benchmarks/results/longbench-v2-codeqa-20-2026-07-17/, PR #171) is a disclosed, cost-bounded 20-of-50 stratified subsample of the
Code Repository Understandingdomain, not the full 50-task domain — full-domain cost was estimated at$460-610 (one pilot task alone cost $3.30 on 659K tokens), disproportionate relative to the other #166 benchmark families ($20-30 each).At n=20, droste's own README already discloses the direct-vs-droste comparison is underpowered: "the two-task score difference from direct-sol is the observed result, not evidence of a population-level separation." droste-terra-luna ties direct-terra and trails direct-sol by 10 points at this sample size — a real, current result, but not a statistically meaningful one given how few tasks separate the arms.
What's needed
Run LongBench-v2 CodeQA at a larger sample size (or the full 50-task domain) once funding allows, to get a droste-vs-direct-model comparison with enough statistical power to actually distinguish the arms, rather than the current small, cost-bounded sample.
Estimate
~$460-610 for the full 50-task x 3-arm domain (established during the original pilot). A partial expansion (e.g. 30-35 tasks, weighted toward the "long" stratum) could meaningfully increase statistical power at a fraction of that cost — worth estimating precisely before committing to full-domain.
Not blocking
This does not block droste#166's current publication (now fully complete) — see PR #171 for the existing, honestly-disclosed 20-task result.