Skip to content

fix(influxdb3_wal): skip empty WalPeriods during replay to avoid panic - #27565

Open
TheManishCode wants to merge 1 commit into
influxdata:mainfrom
TheManishCode:fix-empty-wal-period-replay-panic
Open

TheManishCode wants to merge 1 commit into
influxdata:mainfrom
TheManishCode:fix-empty-wal-period-replay-panic

Conversation

@TheManishCode

Copy link
Copy Markdown

Fixes #26237

Problem

Old WAL files written before an empty WriteBatch was skipped can have their min/max timestamps left at the fold sentinels (i64::MAX/i64::MIN). WalPeriod::new() asserts min_time <= max_time, so replaying one of these pre-existing files on startup panics and crashes the server, with no way to get past it short of manually editing/deleting object-store keys.

Root cause

WalObjectStore::replay() calls FlushBuffer::replay_wal_period, which built a WalPeriod::new(...) unconditionally from each replayed file's min/max timestamps. An empty batch leaves those fields at the sentinel values, tripping the assert. The live flush path (flush_buffer_into_contents_and_responses) doesn't hit this because it inserts a no-op before building an empty period via a plain struct literal — so this only affects replay of old, pre-existing files.

Fix

replay_wal_period now takes the raw (wal_file_number, min_timestamp_ns, max_timestamp_ns) instead of a pre-built WalPeriod, so the guard lives in the one shared function every caller routes through. It still always advances wal_buffer.wal_file_sequence_number, but skips constructing/registering a WalPeriod when the sentinel case is detected, logging a warn! instead. All other replay side effects (file_notifier.notify, snapshot cleanup) are unchanged.

Tests

Added test_replay_wal_period_skips_empty_sentinel_values in influxdb3_wal/src/object_store/tests.rs, asserting no panic on the sentinel case, that snapshot_tracker.num_wal_periods() stays at 0 for it, and that the wal file sequence number still advances; also checks a normal period is still tracked as before.

cargo check -p influxdb3_wal              # Finished, no errors
cargo test -p influxdb3_wal --lib object_store   # 7 passed
cargo test -p influxdb3_wal                # 17 passed, 0 failed

influxdata#26237)

Old WAL files written before an empty WriteBatch was skipped can have
their min/max timestamps left at the fold sentinels (i64::MAX/i64::MIN).
WalPeriod::new asserts min_time <= max_time, so replaying one of these
pre-existing files on startup panicked and crashed the server.

Guard replay_wal_period so it detects the sentinel case and skips
registering a WalPeriod for that file (there's no data in it to track
anyway), while still advancing the wal file sequence number and letting
the rest of the replay path run as before. The live flush path already
avoids this because it writes a no-op before building an empty period,
so this only affects replay of old files.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Skip WalPeriods that are empty

1 participant