"Managed Service for Apache Spark" is the new name for the product formerly known as "Dataproc on Compute Engine" (cluster deployment) and "Google Cloud Serverless for Apache Spark" (serverless deployment).
This tutorial provides information on the availability of the pre-installed connector,
and shows you how make a specific connector version available to Spark
jobs. Example code shows you how to use the Spark BigQuery connector
within a Spark application.
Use the pre-installed connector
The Spark BigQuery connector is pre-installed on and is available to
Spark jobs run on Managed Service for Apache Spark clusters created with image versions
2.1 and later. The pre-installed connector version is listed on the
image version release pages.
Make a specific connector version available to Spark jobs
If you want to use a connector version that is different from a pre-installed
version on a 2.1 or later image version cluster, or if you want to install
the connector on a pre-2.1 image version cluster, follow the instructions in
this section.
gcloud dataproc clusters create CLUSTER_NAME \
--region=REGION \
--image-version=2.2 \
--metadata=SPARK_BQ_CONNECTOR_VERSION or SPARK_BQ_CONNECTOR_URL\
other flags
Notes:
SPARK_BQ_CONNECTOR_VERSION: Specify a connector version.
Spark BigQuery connector versions are listed on the
spark-bigquery-connector/releases
page in GitHub.
Example:
--metadata=SPARK_BQ_CONNECTOR_VERSION=0.42.1
SPARK_BQ_CONNECTOR_URL: Specify a URL that points to the jar in Cloud Storage.
You can specify the URL of a connector listed in the link column in the
Downloading and Using the Connector
in GitHub or the path to a Cloud Storage location where you have
placed a custom connector jar.
Provide the connector jar URL when you submit your job to the cluster using
the Google Cloud console, gcloud CLI, or the Managed Service for Apache Spark
API.
Console
Use the Spark job Jars files item on the Managed Service for Apache Spark
Submit a job page.
The connector writes the data to BigQuery by
first buffering all the data into a Cloud Storage temporary table. Then it
copies all data from into BigQuery in one operation. The
connector attempts to delete the temporary files once the BigQuery
load operation has succeeded and once again when the Spark application terminates.
If the job fails, remove any remaining temporary
Cloud Storage files. Typically, temporary BigQuery
files are located in gs://[bucket]/.spark-bigquery-[jobid]-[UUID].
Configure billing
By default, the project associated with the credentials or service account is
billed for API usage. To bill a different project, set the following
configuration: spark.conf.set("parentProject", "<BILLED-GCP-PROJECT>").
It can also be added to a read or write operation, as follows:
.option("parentProject", "<BILLED-GCP-PROJECT>").
Run the code
Before running this example, create a dataset named "wordcount_dataset" or
change the output dataset in the code to an existing BigQuery dataset in your
Google Cloud project.
Use the
bq command to create
the wordcount_dataset:
bqmkwordcount_dataset
Use the Google Cloud CLI command
to create a Cloud Storage bucket, which will be used to export to
BigQuery:
gcloudstoragebucketscreategs://[bucket]
Scala
Examine the code and replace the [bucket] placeholder with
the Cloud Storage bucket you created earlier.
/* * Remove comment if you are not running in spark-shell. *import org.apache.spark.sql.SparkSessionval spark = SparkSession.builder() .appName("spark-bigquery-demo") .getOrCreate()*/// Use the Cloud Storage bucket for temporary BigQuery export data used// by the connector.valbucket="[bucket]"spark.conf.set("temporaryGcsBucket",bucket)// Load data in from BigQuery. See// https://github.com/GoogleCloudDataproc/spark-bigquery-connector/tree/0.17.3#properties// for option information.valwordsDF=spark.read.bigquery("bigquery-public-data:samples.shakespeare").cache()wordsDF.createOrReplaceTempView("words")// Perform word count.valwordCountDF=spark.sql("SELECT word, SUM(word_count) AS word_count FROM words GROUP BY word")wordCountDF.show()wordCountDF.printSchema()// Saving the data to BigQuery.(wordCountDF.write.format("bigquery").save("wordcount_dataset.wordcount_output"))
Run the code on your cluster
Use SSH to connect to the Managed Service for Apache Spark cluster
master node
On the >Cluster details page, select the VM Instances tab. Then, click
SSH to the right of the name of the cluster master node>
A browser window opens at your home directory on the master node
Create wordcount.scala with the pre-installed vi,
vim, or nano text editor, then paste in the Scala
code from the
Scala code listing
nano wordcount.scala
Launch the spark-shell REPL.
$ spark-shell --jars=gs://spark-lib/bigquery/spark-bigquery-latest.jar
...
Using Scala version ...
Type in expressions to have them evaluated.
Type :help for more information.
...
Spark context available as sc.
...
SQL context available as sqlContext.
scala>
Run wordcount.scala with the :load wordcount.scala command
to create the BigQuery wordcount_output table. The output
listing displays 20 lines
from the wordcount output.
To preview the output table, open the
BigQuery
page, select the wordcount_output table, and then click
Preview.
PySpark
Examine the code and replace the [bucket] placeholder with
the Cloud Storage bucket you created earlier.
#!/usr/bin/env python"""BigQuery I/O PySpark example."""frompyspark.sqlimportSparkSessionspark=SparkSession \
.builder \
.master('yarn') \
.appName('spark-bigquery-demo') \
.getOrCreate()# Use the Cloud Storage bucket for temporary BigQuery export data used# by the connector.bucket="[bucket]"spark.conf.set('temporaryGcsBucket',bucket)# Load data from BigQuery.words=spark.read.format('bigquery') \
.load('bigquery-public-data:samples.shakespeare') \
words.createOrReplaceTempView('words')# Perform word count.word_count=spark.sql('SELECT word, SUM(word_count) AS word_count FROM words GROUP BY word')word_count.show()word_count.printSchema()# Save the data to BigQueryword_count.write.format('bigquery') \
.save('wordcount_dataset.wordcount_output')
Run the code on your cluster
Use SSH to connect to the Managed Service for Apache Spark cluster master node
On the Cluster details page, select the VM Instances tab. Then, click
SSH to the right of the name of the cluster master node
A browser window opens at your home directory on the master node
[[["Easy to understand","easyToUnderstand","thumb-up"],["Solved my problem","solvedMyProblem","thumb-up"],["Other","otherUp","thumb-up"]],[["Hard to understand","hardToUnderstand","thumb-down"],["Incorrect information or sample code","incorrectInformationOrSampleCode","thumb-down"],["Missing the information/samples I need","missingTheInformationSamplesINeed","thumb-down"],["Other","otherDown","thumb-down"]],["Last updated 2026-10-01 UTC."],[],[]]