Use the Spark BigQuery connector

The spark-bigquery-connector is used with Apache Spark to read and write data from and to BigQuery. The connector takes advantage of the BigQuery Storage API when reading data from BigQuery.

This tutorial provides information on the availability of the pre-installed connector, and shows you how make a specific connector version available to Spark jobs. Example code shows you how to use the Spark BigQuery connector within a Spark application.

Use the pre-installed connector

The Spark BigQuery connector is pre-installed on and is available to Spark jobs run on Managed Service for Apache Spark clusters created with image versions 2.1 and later. The pre-installed connector version is listed on the image version release pages.

Make a specific connector version available to Spark jobs

If you want to use a connector version that is different from a pre-installed version on a 2.1 or later image version cluster, or if you want to install the connector on a pre-2.1 image version cluster, follow the instructions in this section.

Important: The spark-bigquery-connector version must be compatible with the Managed Service for Apache Spark cluster image version. See the Connector to Managed Service for Apache Spark Image Compatibility Matrix.

2.1 and later image version clusters

When you create a Managed Service for Apache Spark cluster with a 2.1 or later image version, specify the connector version as cluster metadata.

gcloud CLI example:

gcloud dataproc clusters create CLUSTER_NAME \
    --region=REGION \
    --image-version=2.2 \
    --metadata=SPARK_BQ_CONNECTOR_VERSION or SPARK_BQ_CONNECTOR_URL\
    other flags

Notes:

  • SPARK_BQ_CONNECTOR_VERSION: Specify a connector version. Spark BigQuery connector versions are listed on the spark-bigquery-connector/releases page in GitHub.

    Example:

    --metadata=SPARK_BQ_CONNECTOR_VERSION=0.42.1
    
  • SPARK_BQ_CONNECTOR_URL: Specify a URL that points to the jar in Cloud Storage. You can specify the URL of a connector listed in the link column in the Downloading and Using the Connector in GitHub or the path to a Cloud Storage location where you have placed a custom connector jar.

    Examples:

    --metadata=SPARK_BQ_CONNECTOR_URL=gs://spark-lib/bigquery/spark-3.5-bigquery-0.42.1.jar
    --metadata=SPARK_BQ_CONNECTOR_URL=gs://PATH_TO_CUSTOM_JAR
    

2.0 and earlier image version clusters

You can make the Spark BigQuery connector available to your application in one of the following ways:

  1. Install the spark-bigquery-connector in the Spark jars directory of every node by using the Managed Service for Apache Spark connectors initialization action when you create your cluster.

  2. Provide the connector jar URL when you submit your job to the cluster using the Google Cloud console, gcloud CLI, or the Managed Service for Apache Spark API.

    Console

    Use the Spark job Jars files item on the Managed Service for Apache Spark Submit a job page.