Skip to content

Document completed human ground-truth audit and synchronized results#74

Merged
DavidBakerEffendi merged 4 commits into
mainfrom
58-complete-human-ground-truth-audit-for-remaining-usagebench-languages
Jul 24, 2026
Merged

Document completed human ground-truth audit and synchronized results#74
DavidBakerEffendi merged 4 commits into
mainfrom
58-complete-human-ground-truth-audit-for-remaining-usagebench-languages

Conversation

@DavidBakerEffendi

@DavidBakerEffendi DavidBakerEffendi commented Jul 24, 2026

Copy link
Copy Markdown
Collaborator

Description

Completes issue #58 by documenting the first human review of all 158 current cases and publishing a synchronized exact-range Bifrost-versus-LSP development run.

Key Changes:

  • Add a dedicated audit-status page with coverage, review procedure, trust boundary, and evaluation-promotion path.
  • Present UsageBench as Bifrost's recurring LSP-parity and regression suite, with future competitors supported by the analyzer-neutral contract.
  • Replace the legacy 16 July figures across the homepage, result pages, case comparison, and all language summaries.
  • Report the shared 131-case matrix: 84 exact for both, 32 exact only for Bifrost, 11 exact only for the LSP, and 4 exact for neither.
  • Preserve the development/evaluation boundary: the corpus remains analyzer-informed and legacy_unattributed pending a second independent review, preregistration, and freeze.

Touch Points:

  • docs/src/content/docs/ground-truth-review.md
  • docs/src/content/docs/{index.mdx,overview.md,methodology.md,reproduce.md}
  • docs/src/content/docs/results/{index.md,case-comparison.md}
  • docs/src/content/docs/languages/*.md
  • docs/astro.config.mjs
  • README.md
  • benchmarks/README.md
  • benchmarks/reviews/2026-07-17-DavidBakerEffendi.md

Synchronized Run:

  • UsageBench: 78e66e5dd3589d4543f1b19a8b3566fa9afd644a
  • Bifrost origin/master: 782522b245fc86e3d39b1cdc0488553a1d262212
  • Bifrost full corpus: 133 exact of 152 scoreable, 3 expected gaps, 16 other non-exact, 2 unsupported, 4 not planned
  • Primary LSP profiles: 95 exact of 131 scoreable, 10 position-unverified, 26 hard, 23 unsupported, 4 not planned
  • Runner errors: 0 across Bifrost and all ten primary LSP profiles

Validation:

  • cargo run -- validate benchmarks/cases (35 documents)
  • RUSTDOC=/Users/dave/.rustup/toolchains/stable-aarch64-apple-darwin/bin/rustdoc cargo test (106 passed; doctests passed)
  • npm run check
  • npm run build (18 pages)
  • Rendered inspection of the homepage, synchronized result, and case-comparison pages

Closes #58

@DavidBakerEffendi DavidBakerEffendi changed the title Document completed human ground-truth audit Document completed human ground-truth audit and synchronized results Jul 24, 2026
@DavidBakerEffendi
DavidBakerEffendi marked this pull request as ready for review July 24, 2026 09:08
@DavidBakerEffendi
DavidBakerEffendi merged commit 3470833 into main Jul 24, 2026
3 checks passed
@DavidBakerEffendi
DavidBakerEffendi deleted the 58-complete-human-ground-truth-audit-for-remaining-usagebench-languages branch July 24, 2026 09:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Complete human ground-truth audit for remaining UsageBench languages

1 participant