Skip to content

Wave 1A: canonical PDF IR and ResearchQA benchmark transition - #4

Open
ltczding-gif wants to merge 72 commits into
mainfrom
codex/wave1a-canonical-ir
Open

Wave 1A: canonical PDF IR and ResearchQA benchmark transition#4
ltczding-gif wants to merge 72 commits into
mainfrom
codex/wave1a-canonical-ir

Conversation

@ltczding-gif

@ltczding-gif ltczding-gif commented Jul 27, 2026

Copy link
Copy Markdown
Owner

Summary

  • add the Wave 1A page/span canonical PDF IR with stable extractor/chunker fingerprints and explicit chunk adjacency
  • adopt pinned ResearchQA as the only active optimization benchmark while shelving, not deleting, the unfinished S5 query/gold track
  • add deterministic nested rq-2 / rq-5 / rq-10 / rq-all paper and question indexes across all ten measured domains
  • freeze the ResearchQA quality baseline on canonical C0 chunks, Ollama qwen3-embedding:4b, cosine dense retrieval, and no reranker
  • specify the accepted note-first rq-2 overnight sweep: seven PDF chunkers, four note chunkers, dense/BM25/hybrid retrieval, five source-composition strategies, four reranking depths, and at most sixteen Top-2 confirmation combinations
  • define multi-format SI provenance and native citation coordinates for PDF, DOCX, XLSX, and CSV while keeping external SI outside the ResearchQA gold evidence universe
  • document the external-data/CC-BY-NC boundary and keep downloaded JSONL, papers, SI, generated notes, embeddings, indexes, and raw results out of Git

Validation

  • python -m pytest -q -> 260 passed, 2 skipped
  • pinned ResearchQA source verified: 15,921,446 bytes, SHA-256 681af78bcb1b60d7740a481a9d37ef3af7d9326a72174dc6798c7e87aaa99b73
  • deterministic tiers built offline: rq-2 20 papers/254 questions; rq-5 50/638; rq-10 100/1,263; rq-all 494/6,211
  • live ignored-cache source audit: all 20 rq-2 benchmark PDFs downloaded and parsed (662 pages); official scientific SI found and validated for eight papers (12 files, including PDF/DOCX/XLSX/CSV), with one additional bundled-SI paper
  • design self-review passed: no placeholders or unresolved scope decisions; staged metrics, failure gates, resume contract, model revisions, and stop conditions are explicit

Scope boundary

This PR defines the evaluation substrate and the accepted overnight design. It does not yet implement or run the full strategy sweep, redistribute ResearchQA or paper assets, generate answer-model scores, claim rq-2 quality as a product result, or advance automatically to rq-5.

@ltczding-gif ltczding-gif changed the title Wave 1A: build canonical PDF provenance IR Wave 1A: canonical PDF IR and ResearchQA benchmark transition Jul 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant