Skip to content

Fix range aggregation with fractional bounds on integer columns - #3140

Open
breken-ai wants to merge 1 commit into
quickwit-oss:mainfrom
breken-ai:fix/range-agg-fractional-bounds-int-columns
Open

breken-ai wants to merge 1 commit into
quickwit-oss:mainfrom
breken-ai:fix/range-agg-fractional-bounds-int-columns

Conversation

@breken-ai

Copy link
Copy Markdown

Summary

A range aggregation with fractional bounds counts the wrong documents on integer columns (u64, i64, date, and JSON fields whose values are all integers in a segment).

to_u64_range converts the bounds with f64_to_fastfield_u64, which truncates toward zero. So on values 1, 2, 3, 4, the range {"from": 1.5, "to": 3.5} becomes 1..3. It counts 1 and 2 instead of 2 and 3, and the bucket comes back as "1-3" with from: 1.0, to: 3.0.

It also splits buckets on JSON fields. A segment where json.price holds only integers has an i64 column, and a segment with floats has an f64 column. Their bucket keys differ ("1-3" vs "1.5-3.5"), so the merged result has 6 buckets instead of 3, and each one holds only part of the counts.

This is the same class of bug #3074 fixed for range queries.

Fix

  • On integer columns, round both bounds up with ceil(). from is inclusive and to is exclusive, so 1.5 <= v < 3.5 holds for exactly the integers in 2..4. f64 columns are unchanged.
  • Report the requested bounds in the bucket key, from and to, instead of converting the u64 bounds back. Keys are now the same for every column type, so buckets from int and float segments merge. For integral bounds the keys stay the same as before.
  • get_bucket_pos now uses partition_point to pick the last bucket that starts at or before the value. A fractional range with no integer in it (e.g. 0.5..0.7 becomes 1..1) is empty and shares its start with the next bucket. binary_search_by_key doesn't say which of two equal keys it returns.

Tests

Two new tests in bucket/range.rs:

  • range_fraction_bounds_on_integer_columns: 1.5..3.5 on the u64, i64 and f64 fields, a negative range on i64 (-3.5..0.5 counts -1 and 0), and an empty 0.5..0.7 bucket.
  • range_fraction_bounds_on_json_segments_with_int_and_float_columns: a JSON field with one integer segment and one float segment gives 3 buckets with the right counts.

Both fail on main (left: "1-3" right: "1.5-3.5"; left: 6 right: 3 buckets) and pass with the fix. cargo test -p tantivy --lib passes (1298 passed, 8 ignored), and cargo clippy -p tantivy --tests is clean.

This PR was prepared with the help of an AI coding assistant.

On u64/i64/date columns the range bounds were truncated toward zero, so
`{"from": 1.5, "to": 3.5}` counted the values 1 and 2 instead of 2 and 3,
and the bucket was reported as `1-3`. Round both bounds up instead (`from`
is inclusive and `to` exclusive), and report the requested bounds in the
bucket key, `from` and `to`, so that segments where a JSON field has an
integer column and segments where it has a float column produce the same
bucket keys and merge into one bucket.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant