Skip to content

feat: add roaring bitmaps in posing lists - #446

Open
cheb0 wants to merge 4 commits into
mainfrom
329-lid-bitmaps
Open

feat: add roaring bitmaps in posing lists#446
cheb0 wants to merge 4 commits into
mainfrom
329-lid-bitmaps

Conversation

@cheb0

@cheb0 cheb0 commented Jun 24, 2026

Copy link
Copy Markdown
Collaborator

Description


  • I have read and followed all requirements in CONTRIBUTING.md;
  • I used LLM/AI assistance to make this pull request;

@cheb0

cheb0 commented Jun 24, 2026

Copy link
Copy Markdown
Collaborator Author

@seqbenchbot up 336-compaction mixed

@seqbenchbot

seqbenchbot commented Jun 24, 2026

Copy link
Copy Markdown
Collaborator

Nice, @cheb0 <(-^,^-)=b!

Your request was successfully served.
Identificator for your ongoing benchmark - ec638c8c.

Here is a list of helpful links:

  • Take a look at Grafana dashboard;
  • Live-tailing logs are also available;

Have a great time!

@cheb0

cheb0 commented Jun 24, 2026

Copy link
Copy Markdown
Collaborator Author

@seqbenchbot down ec638c8c

@seqbenchbot

seqbenchbot commented Jun 24, 2026

Copy link
Copy Markdown
Collaborator

Nice, @cheb0 <(-^,^-)=b!

The benchmark with identificator ec638c8c was finished.
I've prepared a summary for you. Click on Show summary button to see it:

Show summary
Query Type mean (ms) stddev (ms) p(50) (ms) p(95) (ms) p(99) (ms) iterations
base comp diff base comp diff base comp diff base comp diff base comp diff base comp diff
bulk
warm 82.28 83.49 +1.47% 33.75 34.08 +0.96% 74.00 74.00 0.00% 149.50 153.00 +2.34% 208.00 206.00 -0.96% 51833.00 51753.00 -0.15%
service:payment-backend-eu
AND k8s_namespace:prod
AND level:[0 to 3]
AND (
    message:'failed'
    OR message:'timeout'
)
warm 82.15 83.20 +1.27% 30.96 31.69 +2.33% 73.00 75.00 +2.74% 143.00 144.00 +0.70% 193.00 194.00 +0.52% 11002.00 10702.00 -2.73%

Have a great time!

@cheb0
cheb0 force-pushed the 329-lid-bitmaps branch from d975fa6 to 206b3a2 Compare June 26, 2026 06:02
@eguguchkin
eguguchkin requested review from dkharms and eguguchkin June 29, 2026 10:59
Base automatically changed from 336-compaction to main July 7, 2026 10:00
@cheb0
cheb0 force-pushed the 329-lid-bitmaps branch from e069c89 to c12589c Compare July 8, 2026 07:39
@codecov-commenter

codecov-commenter commented Jul 8, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 84.70588% with 78 lines in your changes missing coverage. Please review.
✅ Project coverage is 71.49%. Comparing base (a73114c) to head (b35168b).

Files with missing lines Patch % Lines
frac/sealed/lids/block.go 80.00% 25 Missing and 7 partials ⚠️
frac/sealed/lids/iterator_batched_asc.go 60.41% 12 Missing and 7 partials ⚠️
frac/sealed/lids/iterator_batched_desc.go 70.83% 9 Missing and 5 partials ⚠️
node/batch.go 96.77% 5 Missing ⚠️
frac/sealed/lids/iterator_asc.go 91.89% 1 Missing and 2 partials ⚠️
frac/sealed/lids/iterator_desc.go 91.89% 1 Missing and 2 partials ⚠️
cmd/seq-db/seq-db.go 0.00% 1 Missing ⚠️
frac/sealed/lids/loader.go 50.00% 0 Missing and 1 partial ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #446      +/-   ##
==========================================
+ Coverage   71.33%   71.49%   +0.16%     
==========================================
  Files         233      235       +2     
  Lines       18999    19328     +329     
==========================================
+ Hits        13552    13818     +266     
- Misses       4419     4465      +46     
- Partials     1028     1045      +17     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@cheb0
cheb0 force-pushed the 329-lid-bitmaps branch 3 times, most recently from d3f3922 to 4af844b Compare July 10, 2026 16:12
@github-actions

Copy link
Copy Markdown
Contributor

🔴 Performance Degradation

Some benchmarks have degraded compared to the previous run.
Click on Show table button to see full list of degraded benchmarks.

Show table
Name Previous Current Ratio Verdict
Block_Pack-4 59fdb2 b1aa01
104402.00 ns/op 116156.00 ns/op 1.11 🔴

@eguguchkin eguguchkin modified the milestones: v0.75.0, v0.76.0 Jul 13, 2026
@eguguchkin eguguchkin modified the milestones: v0.76.0, v0.77.0 Jul 20, 2026
@cheb0
cheb0 force-pushed the 329-lid-bitmaps branch from 4af844b to 8c75208 Compare July 28, 2026 07:33
@github-actions

Copy link
Copy Markdown
Contributor

🔴 Performance Degradation

Some benchmarks have degraded compared to the previous run.
Click on Show table button to see full list of degraded benchmarks.

Show table
Name Previous Current Ratio Verdict
AggDeep/size=1000-4 e1ed05 f08f82
5152.00 ns/op 6518.00 ns/op 1.27 🔴
AggDeep/size=10000-4 e1ed05 f08f82
51317.00 ns/op 63240.00 ns/op 1.23 🔴
AggDeep/size=1000000-4 e1ed05 f08f82
5230924.00 ns/op 6590823.00 ns/op 1.26 🔴
AggWide/size=1000-4 e1ed05 f08f82
5168.00 ns/op 6943.00 ns/op 1.34 🔴
AggWide/size=10000-4 e1ed05 f08f82
50505.00 ns/op 62225.00 ns/op 1.23 🔴
AggWide/size=1000000-4 e1ed05 f08f82
529.00 B/op 713.00 B/op 1.35 🔴
5671609.00 ns/op 7338963.00 ns/op 1.29 🔴
And/size=1000-4 e1ed05 f08f82
5.05 ns/op 6.42 ns/op 1.27 🔴
And/size=10000-4 e1ed05 f08f82
5.05 ns/op 6.29 ns/op 1.25 🔴
And/size=1000000-4 e1ed05 f08f82
4.52 ns/op 5.65 ns/op 1.25 🔴
AndTree/size=1000-4 e1ed05 f08f82
4.41 ns/op 5.49 ns/op 1.25 🔴
AndTree/size=10000-4 e1ed05 f08f82
4.41 ns/op 5.87 ns/op 1.33 🔴
AndTree/size=1000000-4 e1ed05 f08f82
4.81 ns/op 5.88 ns/op 1.22 🔴
Bitmask-4 e1ed05 f08f82
223724052.00 ns/op 274446568.00 ns/op 1.23 🔴
Block_Pack-4 e1ed05 f08f82
109010.00 ns/op 148163.00 ns/op 1.36 🔴
Block_Unpack-4 e1ed05 f08f82
92335.00 ns/op 116954.00 ns/op 1.27 🔴
BucketClean-4 e1ed05 f08f82
45283.00 ns/op 56984.00 ns/op 1.26 🔴
Complex/size=1000-4 e1ed05 f08f82
5.05 ns/op 6.80 ns/op 1.35 🔴
Complex/size=10000-4 e1ed05 f08f82
5.05 ns/op 6.16 ns/op 1.22 🔴
Complex/size=1000000-4 e1ed05 f08f82
5.54 ns/op 6.97 ns/op 1.26 🔴
ESBulk-4 e1ed05 f08f82
4767.63 MB/s 3905.80 MB/s 0.82 🔴
67615.00 ns/op 82441.00 ns/op 1.22 🔴
FindSequence_Random/extra-large-4 e1ed05 f08f82
9692.29 MB/s 8521.69 MB/s 0.88 🔴
108262.00 ns/op 123048.00 ns/op 1.14 🔴
FindSequence_Random/large-4 e1ed05 f08f82
15730.57 MB/s 13600.23 MB/s 0.86 🔴
1042.00 ns/op 1205.00 ns/op 1.16 🔴
FindSequence_Random/medium-4 e1ed05 f08f82
9811.02 MB/s 8471.12 MB/s 0.86 🔴
104.60 ns/op 120.90 ns/op 1.16 🔴
FindSequence_Random/small-4 e1ed05 f08f82
6282.24 MB/s 5111.38 MB/s 0.81 🔴
40.80 ns/op 50.08 ns/op 1.23 🔴
FindSequence_Random/tiny-4 e1ed05 f08f82
2537.71 MB/s 1999.87 MB/s 0.79 🔴
25.22 ns/op 32.00 ns/op 1.27 🔴
Indexer-4 e1ed05 f08f82
271028262.00 ns/op 317571984.00 ns/op 1.17 🔴
MergeQPRs_ReusingQPR-4 e1ed05 f08f82
3722025.00 ns/op 4382080.00 ns/op 1.18 🔴
MergeSource-4 e1ed05 f08f82
1649170920.00 ns/op 2006907184.00 ns/op 1.22 🔴
MutexListAppend-4 e1ed05 f08f82
184.95 MB/s 166.38 MB/s 0.90 🔴
NAnd/size=1000-4 e1ed05 f08f82
4.68 ns/op 6.14 ns/op 1.31 🔴
NAnd/size=10000-4 e1ed05 f08f82
4.70 ns/op 5.95 ns/op 1.27 🔴
NAnd/size=1000000-4 e1ed05 f08f82
4.79 ns/op 6.13 ns/op 1.28 🔴
Not/size=1000-4 e1ed05 f08f82
5.03 ns/op 6.41 ns/op 1.28 🔴
Not/size=10000-4 e1ed05 f08f82
5.04 ns/op 6.24 ns/op 1.24 🔴
Not/size=1000000-4 e1ed05 f08f82
5.21 ns/op 6.30 ns/op 1.21 🔴
NotEmpty/size=1000-4 e1ed05 f08f82
5.05 ns/op 6.33 ns/op 1.25 🔴
NotEmpty/size=10000-4 e1ed05 f08f82
5.05 ns/op 6.05 ns/op 1.20 🔴
NotEmpty/size=1000000-4 e1ed05 f08f82
5.04 ns/op 6.59 ns/op 1.31 🔴
Or/size=1000-4 e1ed05 f08f82
4.64 ns/op 5.99 ns/op 1.29 🔴
Or/size=10000-4 e1ed05 f08f82
4.67 ns/op 5.91 ns/op 1.27 🔴
Or/size=1000000-4 e1ed05 f08f82
4.71 ns/op 6.02 ns/op 1.28 🔴
OrTree/size=1000-4 e1ed05 f08f82
4.69 ns/op 5.78 ns/op 1.23 🔴
OrTree/size=10000-4 e1ed05 f08f82
4.68 ns/op 5.83 ns/op 1.25 🔴
OrTree/size=1000000-4 e1ed05 f08f82
5.45 ns/op 7.28 ns/op 1.34 🔴
OrTreeNextGeq/size=1000-4 e1ed05 f08f82
4.68 ns/op 6.19 ns/op 1.32 🔴
OrTreeNextGeq/size=10000-4 e1ed05 f08f82
4.68 ns/op 6.20 ns/op 1.32 🔴
OrTreeNextGeq/size=1000000-4 e1ed05 f08f82
5.26 ns/op 6.54 ns/op 1.24 🔴
ParseESTime/es_stdlib-4 e1ed05 f08f82
198.70 ns/op 255.60 ns/op 1.29 🔴
ParseESTime/handwritten-4 e1ed05 f08f82
44.03 ns/op 53.42 ns/op 1.21 🔴
ParseESTime/rfc3339-4 e1ed05 f08f82
57.38 ns/op 71.03 ns/op 1.24 🔴
ProcessDocuments-4 e1ed05 f08f82
92.83 MB/s 76.82 MB/s 0.83 🔴
4444012.00 ns/op 5362895.00 ns/op 1.21 🔴
Sealing_NoSort-4 e1ed05 f08f82
5317.00 allocs/op 8326.00 allocs/op 1.57 🔴
1194959516.00 ns/op 1472758292.00 ns/op 1.23 🔴
Sealing_WithSort-4 e1ed05 f08f82
5387.00 allocs/op 8401.00 allocs/op 1.56 🔴
2294962601.00 ns/op 2863861719.00 ns/op 1.25 🔴
SeqListAppend-4 e1ed05 f08f82
116.04 MB/s 99.29 MB/s 0.86 🔴
137913962.00 ns/op 161138438.00 ns/op 1.17 🔴
SeqQLParsing-4 e1ed05 f08f82
1266.00 ns/op 1674.00 ns/op 1.32 🔴
SeqQLParsingLong-4 e1ed05 f08f82
10973.00 ns/op 13778.00 ns/op 1.26 🔴

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's gonna be interesting to see how bitmaps will affect sealing latency (and system in overall):

benchstat ~/main.txt ~/bitmaps.txt
goos: linux
goarch: amd64
pkg: github.com/ozontech/seq-db/fracmanager
cpu: 12th Gen Intel(R) Core(TM) i5-12600K
                 │ /home/dkharms/main.txt │     /home/dkharms/bitmaps.txt      │
                 │         sec/op         │   sec/op     vs base               │
Sealing_NoSort-4              459.1m ± 2%   468.2m ± 2%  +1.97% (p=0.005 n=10)

                 │ /home/dkharms/main.txt │      /home/dkharms/bitmaps.txt       │
                 │          B/op          │     B/op      vs base                │
Sealing_NoSort-4             23.05Mi ± 0%   26.06Mi ± 0%  +13.05% (p=0.000 n=10)

                 │ /home/dkharms/main.txt │      /home/dkharms/bitmaps.txt      │
                 │       allocs/op        │  allocs/op   vs base                │
Sealing_NoSort-4              5.322k ± 0%   8.309k ± 0%  +56.11% (p=0.000 n=10)

@cheb0 cheb0 Aug 13, 2026

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So far I see they slow down quite literally everything except places where you actually intersect/union them :)

Comment thread frac/sealed/lids/block.go Outdated
Comment thread frac/sealed/lids/block.go
return b.copyLIDsFromBitmap(t, dst)
}

func (b *Block) copyLIDsFromBitmap(ref int32, buf []uint32) []uint32 {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
func (b *Block) copyLIDsFromBitmap(ref int32, buf []uint32) []uint32 {
func (b *Block) appendLIDsFromBitmapTo(ref int32, buf []uint32) []uint32 {

Comment thread node/batch.go

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As I am getting deeper in this pull request I just have several questions about roaring bitmaps in general:

  • Are they concurrent-safe for read-only use (getting max/min, cardinality)?
  • Are they concurrent-safe for (possibly) mutating operations (like Or(bm1, bm2)) or there are different APIs for performing such operations in-place or returning a new copy instead?

Comment thread frac/sealed/lids/block.go
type Block struct {
LIDs []uint32
Offsets []uint32
types []int32 // determines LID list type: delta-encoded (non-negative value) or bitmap (negative value). nil for delta-encoded blocks

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Have you measured whether materializing types actually buys anything vs. a binary search over bitmapIndexes? Like it's pretty easy to grasp but the liability of performing checks like b.types == nil is error-prone (IMHO).

And this decision is not scalable since you encode the type of block into signedness of integer (so there could be at most two different types). So if we are going to introduce different format for storing lids we would still have to rewrite this logic.

It could be easily rewritten as something like:

func (b *Block) chunkRef(i int) (ChunkType, int) {
      k, found := slices.BinarySearch(b.bitmapIndexes, uint32(i))
      if found {
              return ChunkTypeBitmap, k
      }
      return ChunkTypeDelta, i - k
}

Comment thread config/config.go
// BitmapThreshold specifies minimum number of LIDs in the lid list
// which are serialized as bitmap. LIDs lists with more elements use bitmap encoding,
// while smaller lists use delta encoding.
BitmapThreshold int `config:"bitmap_threshold" default:"65536"`

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we add some kind of validation that BitmapThreshold cannot be larger than BlockSize?

Comment thread frac/sealed/lids/block.go
return size
}

func (p *BlockPacker) Pack(b *UnpackedBlock, dst []byte) []byte {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just curious. Have you tried to measure performance metrics when we include whole posting list for token when it exceeds 65k? Of course it heavily depends on density of lids but still it is quite interesting.

However it might break something if we use block capacity as a divider somewhere (I do not remember)...

Comment thread node/node.go
Comment on lines +17 to +19
NextBatch(need int) LIDBatch
// NextBatchGeq returns next batch (LIDs >= minLID). Returns nil when exhausted.
NextBatchGeq(nextLID LID) LIDBatch
NextBatchGeq(need int, nextLID LID) LIDBatch

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Seems like need argument is not used anywhere. And in #479 it was deleted again.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, I deleted it laster. As part of two PRs need will not be added.

Co-authored-by: Daniil <dkharmsd@gmail.com>
@github-actions

Copy link
Copy Markdown
Contributor

🔴 Performance Degradation

Some benchmarks have degraded compared to the previous run.
Click on Show table button to see full list of degraded benchmarks.

Show table
Name Previous Current Ratio Verdict
Indexer-4 a73114 56ac95
678752766.00 B/op 769872608.00 B/op 1.13 🔴
Sealing_NoSort-4 a73114 56ac95
5335.00 allocs/op 8311.00 allocs/op 1.56 🔴
Sealing_WithSort-4 a73114 56ac95
5389.00 allocs/op 8384.00 allocs/op 1.56 🔴

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants