Skip to content

Curation: 10 new terms (model-card capability categories + eval benchmarks) + ARC-AGI update - #18

Merged
umzcio merged 2 commits into
mainfrom
curate/2026-07-25-model-card-terms
Jul 25, 2026
Merged

Curation: 10 new terms (model-card capability categories + eval benchmarks) + ARC-AGI update#18
umzcio merged 2 commits into
mainfrom
curate/2026-07-25-model-card-terms

Conversation

@umzcio

@umzcio umzcio commented Jul 25, 2026

Copy link
Copy Markdown
Owner

Combines the model-card capability→benchmark panel into one PR. (Originally split as #16/#17, but the two sets cross-link each other's slugs and the validator requires every related slug to exist, so they must land together.)

Capability categories (6)

Agentic (Patterns)

  • Agentic Coding (agentic-coding) — autonomous end-to-end software development.
  • Agentic Terminal Coding (agentic-terminal-coding) — agentic coding in a live terminal.
  • Agentic Search (agentic-search) — autonomous multi-step web research.

Reasoning/work axes (Evaluation)

  • Knowledge Work (knowledge-work) — professional office deliverables.
  • Novel Problem-Solving (novel-problem-solving) — solving unfamiliar problems by reasoning, not recall.
  • Multidisciplinary Reasoning (multidisciplinary-reasoning) — accurate reasoning across many fields.

Benchmarks (4, Evaluation)

  • BrowseComp (browsecomp) — OpenAI web-research benchmark → agentic search.
  • GDPval (gdpval) — OpenAI economically-valuable-work benchmark → knowledge work.
  • Terminal-Bench (terminal-bench) — command-line agent tasks; predecessor to Frontier-Bench.
  • Frontier-Bench (frontier-bench) — continuous agent benchmark (74 tasks, finance/biology/hardware beyond coding, top ~34%), by the Terminal-Bench/Harbor team.

Freshness fix (1)

  • ARC-AGI (arc-agi) — noted the newer ARC-AGI-2 and interactive ARC-AGI-3 (still far from solved). No separate v3 entry, to avoid duplication.

Notes

  • computer-use, humanitys-last-exam, osworld, swe-bench, tau-bench from the panel already existed.
  • Capability terms cross-link to the benchmark that measures them, and vice versa.
  • All verified against primary sources (GitHub harbor-framework/frontier-bench, Snorkel leaderboard, ARC Prize, OpenAI). Frontier-Bench details corroborated via the maintainers' own announcements.

All files pass scripts/validate.mjs (959 files valid). Nothing is live until you merge.

umzcio and others added 2 commits July 24, 2026 20:09
…oding/search, knowledge work, novel problem-solving, multidisciplinary reasoning)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ch) + ARC-AGI freshness update

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@umzcio
umzcio merged commit 42341cb into main Jul 25, 2026
1 check passed
@umzcio
umzcio deleted the curate/2026-07-25-model-card-terms branch July 25, 2026 02:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant