Repository navigation
Conversation
On u64/i64/date columns the range bounds were truncated toward zero, so
`{"from": 1.5, "to": 3.5}` counted the values 1 and 2 instead of 2 and 3,
and the bucket was reported as `1-3`. Round both bounds up instead (`from`
is inclusive and `to` exclusive), and report the requested bounds in the
bucket key, `from` and `to`, so that segments where a JSON field has an
integer column and segments where it has a float column produce the same
bucket keys and merge into one bucket.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
A range aggregation with fractional bounds counts the wrong documents on integer columns (u64, i64, date, and JSON fields whose values are all integers in a segment).
to_u64_rangeconverts the bounds withf64_to_fastfield_u64, which truncates toward zero. So on values1, 2, 3, 4, the range{"from": 1.5, "to": 3.5}becomes1..3. It counts1and2instead of2and3, and the bucket comes back as"1-3"withfrom: 1.0, to: 3.0.It also splits buckets on JSON fields. A segment where
json.priceholds only integers has an i64 column, and a segment with floats has an f64 column. Their bucket keys differ ("1-3"vs"1.5-3.5"), so the merged result has 6 buckets instead of 3, and each one holds only part of the counts.This is the same class of bug #3074 fixed for range queries.
Fix
ceil().fromis inclusive andtois exclusive, so1.5 <= v < 3.5holds for exactly the integers in2..4. f64 columns are unchanged.fromandto, instead of converting the u64 bounds back. Keys are now the same for every column type, so buckets from int and float segments merge. For integral bounds the keys stay the same as before.get_bucket_posnow usespartition_pointto pick the last bucket that starts at or before the value. A fractional range with no integer in it (e.g.0.5..0.7becomes1..1) is empty and shares its start with the next bucket.binary_search_by_keydoesn't say which of two equal keys it returns.Tests
Two new tests in
bucket/range.rs:range_fraction_bounds_on_integer_columns:1.5..3.5on the u64, i64 and f64 fields, a negative range on i64 (-3.5..0.5counts-1and0), and an empty0.5..0.7bucket.range_fraction_bounds_on_json_segments_with_int_and_float_columns: a JSON field with one integer segment and one float segment gives 3 buckets with the right counts.Both fail on
main(left: "1-3" right: "1.5-3.5";left: 6 right: 3buckets) and pass with the fix.cargo test -p tantivy --libpasses (1298 passed, 8 ignored), andcargo clippy -p tantivy --testsis clean.This PR was prepared with the help of an AI coding assistant.