Skip to content

fix: skip null and NaN line x instead of failing, and stop zooms and selections past an integer dtype's range from returning nothing - #152

Open
jvdd wants to merge 10 commits into
mainfrom
fix/line-null-nan-x
Open

jvdd wants to merge 10 commits into
mainfrom
fix/line-null-nan-x

Conversation

@jvdd

@jvdd jvdd commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

An ungrouped line on an in-memory frame with a null or NaN in x returned 500 for every request. The line now skips those rows, as histograms do.

  • check_line_x accepts nulls at the start and NaN at the end of a sorted x (Polars sort order), and warns once with their counts. The check stays one collect.
  • The kernel reads only the valid rows through a zero-copy slice, with no filter pass.
  • A null or NaN in the middle of x, or nulls at the end (for example after sort(nulls_last=True), which keeps a sorted flag), still raises "not sorted".
  • The file scan path already dropped these rows. Its output equals the in-memory output.

Integer bounds past the dtype range:

  • A bound past the range of an integer column cast to null. On a line, a pan below 0 on a UInt64 x started the window in the null rows, and a pan above 255 on UInt8 emptied the line, also on data without nulls. A zoomed histogram or 2D histogram on such a column, and a range selection on it, returned nothing.
  • _typed_range_bounds now clamps a numeric integer bound to the dtype range, and a range wholly outside it is empty. The line windows, the histogram grids and the selection predicates all use it.
  • A bound that is not a number (a string or null from a hand-made request) keeps the plain cast it had before.

Measured on a resident 50M-row frame (4 threads, Float64, Float32, Int64 and Datetime x): the data check and the warm requests take the same time as before within noise, the peak memory stays at 0.1 MB (the slice is zero-copy), and a frame with a null prefix and a NaN suffix costs the same as a clean one.

Tests: minmax, lttb and fpcs with a null prefix and a NaN suffix equal the clean frame; scan equals in-memory; a multi-chunk null prefix; one warning per source and column; the sorted-flag cases; line zooms past the dtype range; histogram zooms and range selections past the dtype range equal the Int64 and Float64 results; non-numeric bounds.

Closes #25

@codspeed

codspeed Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Merging this PR will not alter performance

✅ 32 untouched benchmarks


Comparing fix/line-null-nan-x (cdfe9b9) with main (c449485)

Open in CodSpeed

@jvdd jvdd changed the title fix(line): skip null and NaN x values instead of failing with a 500 fix: skip null and NaN line x instead of failing, and stop zooms and selections past an integer dtype's range from returning nothing Oct 6, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Line: skip null and NaN x rows instead of rejecting them

1 participant