Skip to content

[Java] ParquetIO: resolve file records against the read schema - #40344

Open
mxtymoshyk wants to merge 1 commit into
apache:masterfrom
mxtymoshyk:parquetio-read-schema-27234
Open

mxtymoshyk wants to merge 1 commit into
apache:masterfrom
mxtymoshyk:parquetio-read-schema-27234

Conversation

@mxtymoshyk

Copy link
Copy Markdown
Contributor

Fixes #27234

ParquetIO.read(schema) and ParquetIO.readFiles(schema) fail when the schema has a field the file doesn't have, which is the usual case after you add a field to a class and read older files. This PR makes ParquetIO use that schema as the Avro reader schema, so parquet-avro resolves each file record against it by field name.

Reproduction

Write a file with two fields, then read it with a schema that adds a nullable third field:

Schema oldSchema = /* {firstName: string, middleName: string} */;
Schema newSchema = /* same, plus {"name": "lastName", "type": ["null", "string"], "default": null} */;

// write with oldSchema via FileIO.write().via(ParquetIO.sink(oldSchema)) ...

pipeline.apply(ParquetIO.read(newSchema).from(path));
pipeline.run().waitUntilFinish();

On current master this fails with:

org.apache.parquet.io.ParquetDecodingException: Can not read value at 1 in block 1 in file ...
Caused by: java.lang.ArrayIndexOutOfBoundsException: Index 2 out of bounds for length 2
    at org.apache.avro.generic.GenericData$Record.get(GenericData.java:289)
    ...
    at org.apache.avro.generic.GenericDatumWriter.writeRecord(GenericDatumWriter.java:234)

A related case fails without an exception. If the read schema leaves out a column that isn't the last one (for example, reading {name, id} files with {id}), each record comes back with the name value in the id field.

Root cause

ReadFiles only used the schema to build the output coder (AvroCoder or a Beam SchemaCoder). SplitReadFn never got it, and nothing called AvroReadSupport.setAvroReadSchema. So parquet-avro built each GenericRecord with the writer schema stored in the file footer, and AvroCoder then encoded it by field position against the user's schema. When the two schemas differ in length or order, GenericDatumWriter reads the wrong positions or runs past the end of the record.

Workaround for released versions

Pass the read schema through the Hadoop configuration:

ParquetIO.read(newSchema)
    .from(path)
    .withConfiguration(
        Collections.singletonMap("parquet.avro.read.schema", newSchema.toString()));

This covers added fields. It won't help with a read schema that has fewer columns than the file, because parquet-avro rejects file columns missing from the read schema (Parquet/Avro schema mismatch: Avro field 'x' not found). Use withProjection for that case.

Changes

ReadFiles passes the output schema to SplitReadFn. That's the schema passed to readFiles()/read(), or the encoder schema when withProjection is set, the same one the coder uses. SplitReadFn sets it as parquet.avro.read.schema unless the user already put that key in the configuration.

Without a projection, SplitReadFn also drops top-level file columns that match no field name or alias in the read schema. parquet-avro requires every requested column to exist in the read schema, so without this step a subset schema that works today by accident (a prefix of the file's columns) would start failing.

parseGenericRecords and parseFilesGenericRecords take no schema and are unchanged.

Behavior change

Fields now resolve by name (or Avro alias) instead of by position. A pipeline that reads files whose column names differ from its Avro field names, for example files written by another tool and read with a hand-written schema, used to get position-matched data and will now fail with Avro field 'x' not found or with a null in a non-nullable field. I think that failure is correct, since positional matching gives the wrong data as soon as column order differs. Reviewers may want this listed under Breaking Changes in CHANGES.md instead of Bugfixes. I'm happy to move it.

Tests

New tests in ParquetIOTest:

  • testReadWithAddedNullableField: the reproduction above. The new field reads as null.
  • testReadFilesWithAddedFieldWithDefault: a non-null default ("unknown") gets filled in, which shows real schema resolution and not only null padding. Goes through readFiles().
  • testReadWithSubsetSchemaWithoutProjection: a schema that drops the first column returns the right ids.
  • testReadWithProjectionAndAddedField: withProjection with an encoder schema that has an extra field.
  • testReadSchemaFromConfigurationIsNotOverridden: a parquet.avro.read.schema set in the configuration wins.
  • testPruneToReadSchema: unit test for the column pruning, including alias matching and the no-match case.

With the fix disabled, the first four fail: three with the ParquetDecodingException shown above, and the subset test with wrong values. The existing 19 tests pass unchanged. ./gradlew :sdks:java:io:parquet:check :sdks:java:io:parquet:javadoc passes locally.

Not covered

Pruning only works on top-level columns. A nested record in the read schema that has fewer fields than the file's nested group still fails in parquet-avro, the same as on master.


  • Mention the appropriate issue in your description (for example: addresses #123), if applicable. This will automatically add a link to the pull request in the issue. If you would like the issue to automatically close on merging the pull request, comment fixes #<ISSUE NUMBER> instead.
  • Update CHANGES.md with noteworthy changes.
  • If this contribution is large, please file an Apache Individual Contributor License Agreement.

See the Contributor Guide for more tips on how to make review process smoother.

@github-actions

Copy link
Copy Markdown
Contributor

Assigning reviewers:

R: @Abacn for label java.

Note: If you would like to opt out of this review, comment assign to next reviewer.

Available commands:

  • stop reviewer notifications - opt out of the automated review tooling
  • remind me after tests pass - tag the comment author after tests pass
  • waiting on author - shift the attention set back to the author (any comment or push by the author will return the attention set to the reviewers)

The PR bot will only process comments in the main thread (not review comments).

ParquetIO.read(schema) and readFiles(schema) used the schema only for
the output coder. parquet-avro built records with the file's writer
schema, so AvroCoder encoded them by position against a schema with a
different shape. Reading files that lack a newly added field failed
with ArrayIndexOutOfBoundsException, and a schema that omits a leading
column returned values from the wrong column.

SplitReadFn now sets parquet.avro.read.schema to the output schema
(the encoder schema when a projection is set), unless the user already
set it. Without a projection, it also drops file columns that the read
schema has no field for, since the Avro converter rejects them.

Fixes apache#27234
@mxtymoshyk
mxtymoshyk force-pushed the parquetio-read-schema-27234 branch from 29b05c2 to fcd6608 Compare October 3, 2026 06:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: ParquetIO failing to read with specified schema

1 participant