[Java] ParquetIO: resolve file records against the read schema - #40344
Open
mxtymoshyk wants to merge 1 commit into
Open
mxtymoshyk wants to merge 1 commit into
mxtymoshyk wants to merge 1 commit into
Conversation
Contributor
|
Assigning reviewers: R: @Abacn for label java. Note: If you would like to opt out of this review, comment Available commands:
The PR bot will only process comments in the main thread (not review comments). |
mxtymoshyk
force-pushed
the
parquetio-read-schema-27234
branch
from
October 1, 2026 14:04
d6bb15c to
29b05c2
Compare
ParquetIO.read(schema) and readFiles(schema) used the schema only for the output coder. parquet-avro built records with the file's writer schema, so AvroCoder encoded them by position against a schema with a different shape. Reading files that lack a newly added field failed with ArrayIndexOutOfBoundsException, and a schema that omits a leading column returned values from the wrong column. SplitReadFn now sets parquet.avro.read.schema to the output schema (the encoder schema when a projection is set), unless the user already set it. Without a projection, it also drops file columns that the read schema has no field for, since the Avro converter rejects them. Fixes apache#27234
mxtymoshyk
force-pushed
the
parquetio-read-schema-27234
branch
from
October 3, 2026 06:03
29b05c2 to
fcd6608
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #27234
ParquetIO.read(schema)andParquetIO.readFiles(schema)fail when the schema has a field the file doesn't have, which is the usual case after you add a field to a class and read older files. This PR makes ParquetIO use that schema as the Avro reader schema, so parquet-avro resolves each file record against it by field name.Reproduction
Write a file with two fields, then read it with a schema that adds a nullable third field:
On current master this fails with:
A related case fails without an exception. If the read schema leaves out a column that isn't the last one (for example, reading
{name, id}files with{id}), each record comes back with thenamevalue in theidfield.Root cause
ReadFilesonly used the schema to build the output coder (AvroCoderor a BeamSchemaCoder).SplitReadFnnever got it, and nothing calledAvroReadSupport.setAvroReadSchema. So parquet-avro built eachGenericRecordwith the writer schema stored in the file footer, andAvroCoderthen encoded it by field position against the user's schema. When the two schemas differ in length or order,GenericDatumWriterreads the wrong positions or runs past the end of the record.Workaround for released versions
Pass the read schema through the Hadoop configuration:
This covers added fields. It won't help with a read schema that has fewer columns than the file, because parquet-avro rejects file columns missing from the read schema (
Parquet/Avro schema mismatch: Avro field 'x' not found). UsewithProjectionfor that case.Changes
ReadFilespasses the output schema toSplitReadFn. That's the schema passed toreadFiles()/read(), or the encoder schema whenwithProjectionis set, the same one the coder uses.SplitReadFnsets it asparquet.avro.read.schemaunless the user already put that key in the configuration.Without a projection,
SplitReadFnalso drops top-level file columns that match no field name or alias in the read schema. parquet-avro requires every requested column to exist in the read schema, so without this step a subset schema that works today by accident (a prefix of the file's columns) would start failing.parseGenericRecordsandparseFilesGenericRecordstake no schema and are unchanged.Behavior change
Fields now resolve by name (or Avro alias) instead of by position. A pipeline that reads files whose column names differ from its Avro field names, for example files written by another tool and read with a hand-written schema, used to get position-matched data and will now fail with
Avro field 'x' not foundor with a null in a non-nullable field. I think that failure is correct, since positional matching gives the wrong data as soon as column order differs. Reviewers may want this listed under Breaking Changes inCHANGES.mdinstead of Bugfixes. I'm happy to move it.Tests
New tests in
ParquetIOTest:testReadWithAddedNullableField: the reproduction above. The new field reads as null.testReadFilesWithAddedFieldWithDefault: a non-null default ("unknown") gets filled in, which shows real schema resolution and not only null padding. Goes throughreadFiles().testReadWithSubsetSchemaWithoutProjection: a schema that drops the first column returns the right ids.testReadWithProjectionAndAddedField:withProjectionwith an encoder schema that has an extra field.testReadSchemaFromConfigurationIsNotOverridden: aparquet.avro.read.schemaset in the configuration wins.testPruneToReadSchema: unit test for the column pruning, including alias matching and the no-match case.With the fix disabled, the first four fail: three with the
ParquetDecodingExceptionshown above, and the subset test with wrong values. The existing 19 tests pass unchanged../gradlew :sdks:java:io:parquet:check :sdks:java:io:parquet:javadocpasses locally.Not covered
Pruning only works on top-level columns. A nested record in the read schema that has fewer fields than the file's nested group still fails in parquet-avro, the same as on master.
addresses #123), if applicable. This will automatically add a link to the pull request in the issue. If you would like the issue to automatically close on merging the pull request, commentfixes #<ISSUE NUMBER>instead.CHANGES.mdwith noteworthy changes.See the Contributor Guide for more tips on how to make review process smoother.