Curation: 4 new benchmarks (BrowseComp, GDPval, Terminal-Bench, Frontier-Bench) + ARC-AGI update - #17
Closed
umzcio wants to merge 1 commit into
Closed
Curation: 4 new benchmarks (BrowseComp, GDPval, Terminal-Bench, Frontier-Bench) + ARC-AGI update#17umzcio wants to merge 1 commit into
umzcio wants to merge 1 commit into
Conversation
…ch) + ARC-AGI freshness update Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
New terms (4) — the benchmarks behind the capability axes
browsecomp, Evaluation) — OpenAI's benchmark for hard, multi-step web research (measures agentic search).gdpval, Evaluation) — OpenAI's benchmark for economically valuable professional work across many occupations (measures knowledge work).terminal-bench, Evaluation) — agent tasks inside a command-line terminal; predecessor to Frontier-Bench.frontier-bench, Evaluation) — continuously-updated agent benchmark (74 tasks; finance/biology/music/hardware beyond coding; top agents ~34%), by the Terminal-Bench/Harbor team.Freshness fix (1)
arc-agi) — added a line noting the newer ARC-AGI-2 and the interactive ARC-AGI-3 (explore-a-game-with-no-rules), both still far from solved. Kept slug/structure; no separate ARC-AGI-3 entry to avoid duplication.Sourcing
harbor-framework/frontier-bench), Snorkel AI leaderboard, ARC Prize, OpenAI benchmark pages. Frontier-Bench details (74 tasks, ~34%, Terminal-Bench/Harbor lineage) corroborated via the maintainers' own announcements.All files pass
scripts/validate.mjs(959 files valid). Nothing is live until you merge.Sources: Frontier-Bench (GitHub), Snorkel AI — Frontier-Bench, ARC Prize, DataCamp — ARC-AGI-3