Goal
Add LongBench-Pro as a separately reported public diagnostic for general long-context comprehension and RLM transfer. It must remain an independent suite rather than being averaged into unrelated benchmark verdicts.
References:
Requirements
- Verify and pin the immutable dataset revision, configuration/splits, license, attribution, redistribution, and derived-artifact rules.
- Preserve the published English/Chinese, task, full/partial-context, length, and difficulty dimensions.
- Implement the standard task-specific metrics: NDCG, pairwise accuracy, accuracy, F1, SubEM, and the published summarization metrics.
- Keep source questions, answers, rationales, and any restricted fields subject to the dataset license and publication rules.
- Add method-neutral direct and RLM execution adapters with identical answer contracts and failure-inclusive denominators.
- Record exact model routes, reasoning controls, prompt revision, execution limits, pricing provenance, usage, cost, latency, and typed failures.
- Produce immutable manifests, SHA-256 hashes, fresh no-overwrite artifact directories, and offline fixture regeneration.
- Report results by language, task, context dependency, length, and difficulty; never collapse context-limit failures into truncation.
- Clearly label the suite as a diagnostic. Do not combine it with another suite to support a headline claim.
- Run full tests, public-hygiene checks, Claude review, and a concise architecture review before merge.
Non-goals
- Do not tune prompts or execution policy on held-out LongBench-Pro results.
- Do not use LongBench-Pro as a replacement for domain-specific evaluation.
- Do not publish quantitative claims without immutable version-matched artifacts.
Goal
Add LongBench-Pro as a separately reported public diagnostic for general long-context comprehension and RLM transfer. It must remain an independent suite rather than being averaged into unrelated benchmark verdicts.
References:
Requirements
Non-goals