Skip to content

feat(query): push LIMIT down across shard groups for raw selects - #27690

Open
alrieckert wants to merge 1 commit into
influxdata:master-1.xfrom
alrieckert:limit-pushdown
Open

alrieckert wants to merge 1 commit into
influxdata:master-1.xfrom
alrieckert:limit-pushdown

Conversation

@alrieckert

Copy link
Copy Markdown

SELECT ... LIMIT N on a single, ungrouped measurement with no aggregate call currently scans every shard in the query's time range before LIMIT is ever applied: Shards.CreateIterator unconditionally calls CreateIterator on every shard, and the per-series LimitIterator that eventually trims the result never stops pulling from its input early - it only returns once the fully-merged, all-shards input itself is exhausted. So LIMIT reduces rows returned, not work done, which is expensive once a measurement spans many shards/shard groups with no tag filter to narrow the series set.

Add a narrow, provably-safe fast path: when a fetch is for a single, non-regex Measurement, has no GROUP BY (tags or time), no SLIMIT/SOFFSET, and is the sole source feeding the statement's LIMIT/OFFSET (IteratorOptions.GlobalLimitEligible, set only by buildAuxIterator), LocalShardMapping.CreateIterator lazily opens shard groups one at a time - in the order matching ORDER BY's direction, reversed for DESC - and stops opening further groups once enough rows have already been produced to satisfy Limit+Offset (query.NewLazyGroupChainIterator). All point filtering still happens in the existing, unchanged top-level LimitIterator; the new iterator is a pure counting pass-through. Every other query shape (GROUP BY, regex, multi-measurement, SLIMIT/SOFFSET, aggregates) is completely unaffected and takes the original, unchanged path.

  • I've read the contributing section of the project README.
  • Signed CLA (if not already signed).

SELECT ... LIMIT N on a single, ungrouped measurement with no aggregate
call currently scans every shard in the query's time range before LIMIT is
ever applied: Shards.CreateIterator unconditionally calls CreateIterator on
every shard, and the per-series LimitIterator that eventually trims the
result never stops pulling from its input early - it only returns once the
fully-merged, all-shards input itself is exhausted. So LIMIT reduces rows
returned, not work done, which is expensive once a measurement spans many
shards/shard groups with no tag filter to narrow the series set.

Add a narrow, provably-safe fast path: when a fetch is for a single,
non-regex Measurement, has no GROUP BY (tags or time), no SLIMIT/SOFFSET,
and is the sole source feeding the statement's LIMIT/OFFSET
(IteratorOptions.GlobalLimitEligible, set only by buildAuxIterator),
LocalShardMapping.CreateIterator lazily opens shard groups one at a time -
in the order matching ORDER BY's direction, reversed for DESC - and stops
opening further groups once enough rows have already been produced to
satisfy Limit+Offset (query.NewLazyGroupChainIterator). All point
filtering still happens in the existing, unchanged top-level LimitIterator;
the new iterator is a pure counting pass-through. Every other query shape
(GROUP BY, regex, multi-measurement, SLIMIT/SOFFSET, aggregates) is
completely unaffected and takes the original, unchanged path.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant