Add Docling document model, visual enrichment, and processing trace - #344
Open
farhat-is-coding wants to merge 6 commits into
Open
Add Docling document model, visual enrichment, and processing trace#344farhat-is-coding wants to merge 6 commits into
farhat-is-coding wants to merge 6 commits into
Conversation
Tighten comments to one-liners per rules.md, move needs_ocr derivation onto the File model, drop the DOCUMENT_ENRICHMENT_MAX_PAGES cap (silent page truncation), and delete the internal corpus inspection script.
farhat-is-coding
marked this pull request as ready for review
August 15, 2026 23:49
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
Why
The previous upload path flattened PDFs too early, losing useful document structure and leaving figures unsearchable. Failures were also attributed to the overall upload rather than the exact parsing, enrichment, embedding, or indexing stage.
This change preserves Docling's structural model, adds searchable visual meaning with source context, and makes processing behavior, cost, caching, and failures inspectable.
Impact
Mixed text-and-image PDFs now retain headings, tables, captions, and visual descriptions in retrieval text. Textless pages can recover searchable transcription through Gemini. Existing non-Docling and non-visual documents continue through the normal path.
Internal-reference resolution and the user-approved large scanned-document OCR workflow are intentionally out of scope for follow-up PRs.
Validation