Curation: 10 new terms (model-card capability categories + eval benchmarks) + ARC-AGI update - #18
Merged
Merged
Conversation
…oding/search, knowledge work, novel problem-solving, multidisciplinary reasoning) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ch) + ARC-AGI freshness update Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Combines the model-card capability→benchmark panel into one PR. (Originally split as #16/#17, but the two sets cross-link each other's slugs and the validator requires every
relatedslug to exist, so they must land together.)Capability categories (6)
Agentic (Patterns)
agentic-coding) — autonomous end-to-end software development.agentic-terminal-coding) — agentic coding in a live terminal.agentic-search) — autonomous multi-step web research.Reasoning/work axes (Evaluation)
knowledge-work) — professional office deliverables.novel-problem-solving) — solving unfamiliar problems by reasoning, not recall.multidisciplinary-reasoning) — accurate reasoning across many fields.Benchmarks (4, Evaluation)
browsecomp) — OpenAI web-research benchmark → agentic search.gdpval) — OpenAI economically-valuable-work benchmark → knowledge work.terminal-bench) — command-line agent tasks; predecessor to Frontier-Bench.frontier-bench) — continuous agent benchmark (74 tasks, finance/biology/hardware beyond coding, top ~34%), by the Terminal-Bench/Harbor team.Freshness fix (1)
arc-agi) — noted the newer ARC-AGI-2 and interactive ARC-AGI-3 (still far from solved). No separate v3 entry, to avoid duplication.Notes
computer-use,humanitys-last-exam,osworld,swe-bench,tau-benchfrom the panel already existed.harbor-framework/frontier-bench, Snorkel leaderboard, ARC Prize, OpenAI). Frontier-Bench details corroborated via the maintainers' own announcements.All files pass
scripts/validate.mjs(959 files valid). Nothing is live until you merge.