⚡ Optimize crawler database insertions (N+1 fix)#23
Conversation
Optimizes the crawl link processing loop by replacing N+1 database queries with batch operations. Introduces `PageRepository.upsertMany` and `EdgeRepository.insertEdges` to handle bulk updates efficiently. Uses `RETURNING id` in `upsertMany` to resolve page IDs in a single transaction. Chunks `getPagesByUrls` queries to avoid SQLite variable limits. Measured performance improvement: ~3.3x speedup on link processing (from ~8000 ops/sec to ~26000 ops/sec in-memory baseline).
|
👋 Jules, reporting for duty! I'm here to lend a hand with this pull request. When you start a review, I'll add a 👀 emoji to each comment to let you know I've read it. I'll focus on feedback directed at me and will do my best to stay out of conversations between you and other bots or reviewers to keep the noise down. I'll push a commit with your requested changes shortly after. Please note there might be a delay between these steps, but rest assured I'm on the job! For more direct control, you can switch me to Reactive Mode. When this mode is on, I will only act on comments where you specifically mention me with New to Jules? Learn more at jules.google/docs. For security, I will only act on instructions from the user who triggered this task. |
Optimizes the crawler's link processing loop to eliminate N+1 query patterns.
upsertManytoPageRepositorywhich usesINSERT ... ON CONFLICT ... RETURNING idinside a transaction to upsert pages and retrieve their IDs efficiently.insertEdgestoEdgeRepositoryto bulk insert edges in a single transaction.crawl.tsto collect links and process them in batches.IN (...)queries ingetPagesByUrlsto avoid SQLite variable limits.PR created automatically by Jules for task 9562293431074619775 started by @saurabhsharma2u