Scan 2026-07-31: 8 new eval finds - #69
Open
github-actions[bot] wants to merge 1 commit into
Open
Conversation
Seven README entries across §5, §6, §9, §10; one MENTIONS.md entry.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Additions (2026-07-31 scan)
All 8 entries are URL-verified against live pages; verbatim quotes are from the actual sources.
README.md — §5 (Eval infrastructure)
Do Automated Evals Work? — Antaripa Saha & Hamel Husain (Parlance Labs, Jul 11 2026) → §5f
PACE: A Proxy for Agentic Capability Evaluation — Song, Sutawika, Liu et al. (CMU / Berkeley, Jul 2026) → §5
README.md — §6 (Benchmark integrity)
Separating signal from noise in coding evaluations — OpenAI (Jul 8 2026) → §6
Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation — Bhat, Vaghasiya, Mohsin, Aali → §6
README.md — §9 (Agent-specific evaluation)
Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT — METR (Jul 21 2026) → §9
Behavior specs: an open standard for supervising long-horizon agents — Braintrust + Basis (Jul 29 2026) → §9
README.md — §10 (Safety / adversarial evaluation)
Investigating three real-world incidents in our cybersecurity evaluations — Anthropic (Jul 30 2026) → §10
MENTIONS.md
Harness Engineering for Self-Improvement — Lilian Weng (Jul 4 2026) → MENTIONS.md
What was rejected and why