From b4a4f9b30fbaf700e0ee01158e97eb5b63afe9cf Mon Sep 17 00:00:00 2001 From: genitrix Date: Tue, 11 Aug 2026 00:50:59 +0800 Subject: [PATCH] Add Dr. Bench --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 4ea2ad2..817a072 100644 --- a/README.md +++ b/README.md @@ -386,6 +386,7 @@ Most "awesome" lists are link dumps. This one is **annotated and verified**: eve - **[ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery](https://github.com/OSU-NLP-Group/ScienceAgentBench)** โ€” OSU-NLP Group (Ohio State) โ€” ยท *benchmark* โ€” 102 expert-validated tasks from 44 peer-reviewed papers; grades self-contained Python programs by execution + success rate; best agent solves only ~34% (ICLR 2025). ๐Ÿ†• - **[CORE-Bench: Computational Reproducibility Agent Benchmark](https://arxiv.org/abs/2409.11363)** โ€” Siegel, Kapoor, Narayanan et al. (Princeton) โ€” ยท *benchmark* โ€” 270 tasks over 90 papers (CS/social science/medicine) that grade whether an agent can reproduce published results from code+data; from the Princeton AI-Snake-Oil group. - **[DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents](https://arxiv.org/abs/2506.11763)** โ€” Mingxuan Du et al. โ€” ยท *benchmark* โ€” 100 PhD-level tasks across 22 fields; reference-based adaptive-rubric grader for analyst-grade citation-rich reports, validated for human-judgment alignment โ€” the standard deep-research-report eval. ๐Ÿ†• +- **[Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports](https://arxiv.org/abs/2510.02190)** โ€” Yao et al. โ€” ยท *benchmark* โ€” 214 expert-curated deep-research tasks across 10 domains, evaluating long reports for semantic quality, topical focus, and retrieval trustworthiness. ๐Ÿ†• - **[BixBench: A Comprehensive Benchmark for LLM-based Agents in Computational Biology](https://arxiv.org/abs/2503.00096)** โ€” FutureHouse + ScienceMachine โ€” ยท *benchmark* โ€” 50+ real bioinformatics analysis scenarios with ~300 open-answer questions over multi-step Jupyter trajectories; frontier models hit only ~17% โ€” serious wet-lab-adjacent science agent eval. ๐Ÿ†• - **[Introducing LifeSciBench](https://openai.com/index/introducing-life-sci-bench/)** โ€” OpenAI โ€” ยท *benchmark* โ€” ๐Ÿ†• 750 expert-authored life-science research tasks (7 workflows ร— 7 biological domains) graded by 19,020 rubric criteria from 173 PhD-level scientist contributors, independently validated by 453 expert reviewers; requires interpreting genomic sequence files, chemical structures, and experimental figures; best model (GPT-Rosalind) scores 36.1% overall โ€” the largest expert-rubric-graded wet-lab-workflow agent benchmark. - **[Gaia2 and ARE: Scaling Up Agent Environments and Evaluations](https://arxiv.org/abs/2509.17158)** โ€” Meta (Meta Agents Research Environments) โ€” ยท *benchmark* โ€” Successor to GAIA: dynamic, time-driven, multi-agent simulated environments with async world events and a verifiable scenario grader; frontier success ~42% โ€” the serious general-assistant env from Meta. ๐Ÿ†• @@ -573,4 +574,3 @@ To the extent possible under law, [BenchFlow](https://benchflow.ai) and contribu -