feat(writer): cascade RLE like Rust's IntRLEScheme (issue #410) - #437
Merged
Merged
Conversation
Parity with the reference compressor: fastlanes.rle overrides encodeCascade and hands values, indices and offsets to the compressor as open children (Rust ids values=0, indices=1, offsets=2 match our wire order), with Rust's rle_descendant_exclusions: Dict and Sparse barred on indices and offsets. Rust's RunEnd rule there is commented out upstream as unsound, so it is not ported. RLE names itself on every child, matching Rust's "no scheme twice in one chain". The chunked run computation moves into one Runs helper shared by encode, encodeBool and encodeCascade. Output is byte-identical on every size comparison; all interop suites green. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
dfa1
added a commit
that referenced
this pull request
Oct 3, 2026
Parity gap found while porting IntRLEScheme (#437): Rust also RLE-encodes floats, while RleEncodingEncoder refused them. It now accepts every primitive. Runs compare raw bits (as toLongs already did for floats), so the encoding is lossless: -0.0, +0.0 and NaN payloads stay distinct. The cascade's values child is the column's own float[]/double[], so ALP can compete on the run values; Rust's float scheme shares the int one's children and exclusions, so nothing else changes. vortex-jni reads Java-written float RLE back bit for bit, forced and through the cascade. Size comparisons are byte-identical. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
dfa1
added a commit
that referenced
this pull request
Oct 3, 2026
Parity gap found while porting IntRLEScheme (#437): Rust also RLE-encodes floats, while RleEncodingEncoder refused them. It now accepts every primitive. Runs compare raw bits (as toLongs already did for floats), so the encoding is lossless: -0.0, +0.0 and NaN payloads stay distinct. The cascade's values child is the column's own float[]/double[], so ALP can compete on the run values; Rust's float scheme shares the int one's children and exclusions, so nothing else changes. vortex-jni reads Java-written float RLE back bit for bit, forced and through the cascade. Size comparisons are byte-identical. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #410, Rust-parity port number 3 after RunEnd (#435) and ZigZag (#436).
RleEncodingEncoder.encodeCascadehands the threefastlanes.rlechildren to the compressor as open slots instead of raw buffers. Rust's ids (values=0, indices=1, offsets=2) happen to match our wire order:rle_descendant_exclusions)fromLongsArray)U16, per padded rowU64, per chunkis_excluded("no scheme appears twice in any chain"). Our slot sets union down the chain, so the effect is the same as Rust's push rules.Runshelper shared byencode,encodeBoolandencodeCascade, so there are no longer three copies.Parity gap noted, not addressed here: Rust also has a float RLE scheme; ours accepts integers only.
Size: byte-identical on every FileSizeComparison case.
Tests:
RleEncodingEncoderTest.Cascade(values/indices/offsets across 2 FastLanes chunks, child dtypes, exclusions, empty input). Green: writer and reader unit suites, plus 367 integration tests including vortex-jni interop.🤖 Generated with Claude Code