Create and use data profile scans

Knowledge Catalog (formerly Dataplex Universal Catalog) lets you identify common statistical characteristics (common values, data distribution, null counts) of the columns in your BigQuery tables. This information helps you to understand and analyze your data more effectively.

For more information about Knowledge Catalog data profile scans, see About data profiling.

Before you begin

Enable the Dataplex API, if it is not already enabled.

Roles required to enable APIs

To enable APIs, you need the serviceusage.services.enable permission. If you created the project, then you likely already have this permission through the Owner role (roles/owner). Otherwise, you can get this permission through the Service Usage Admin role (roles/serviceusage.serviceUsageAdmin). Learn how to grant roles.

Enable the API

Required roles and permissions

This section describes the IAM roles and permissions needed to use Knowledge Catalog data profile scans.

User roles and permissions

To get the permissions that you need to create and manage data profile scans, ask your administrator to grant you the following IAM roles:

  • Create, run, update, and delete data profile scans: Dataplex DataScan Editor (roles/dataplex.dataScanEditor) on the project containing the data scan
  • View data profile scan results, jobs, and history: Dataplex DataScan Viewer (roles/dataplex.dataScanViewer) on the project containing the data scan
  • Publish data profile scan results to Knowledge Catalog: Dataplex Catalog Editor (roles/dataplex.catalogEditor) on the @bigquery entry group
  • View published data profile scan results in BigQuery on the Data profile tab: BigQuery Data Viewer (roles/bigquery.dataViewer) on the table
  • Run data profile scans:
  • Run data profile scans against BigQuery external tables that use Cloud Storage data:
  • Run data profile scans for Iceberg REST Catalog, SAP BDC Delta Lake, and Apache Hive tables on Google Cloud Lakehouse: BigLake Viewer (roles/biglake.viewer) on the tables being scanned
  • Export data profile scan results to a BigQuery table: BigQuery Data Editor (roles/bigquery.dataEditor) on the table

For more information about granting roles, see Manage access to projects, folders, and organizations.

These predefined roles contain the permissions required to create and manage data profile scans. To see the exact permissions that are required, expand the Required permissions section:

Required permissions

The following permissions are required to create and manage data profile scans:

  • Create, run, update, and delete data profile scans:
    • dataplex.datascans.create on project
    • dataplex.datascans.update on data scan
    • dataplex.datascans.delete on data scan
    • dataplex.datascans.run on data scan
    • dataplex.datascans.get on data scan
    • dataplex.datascans.list on project
    • dataplex.dataScanJobs.get on data scan job
    • dataplex.dataScanJobs.list on data scan
  • View data profile scan results, jobs, and history:
    • dataplex.datascans.getData on data scan
    • dataplex.datascans.list on project
    • dataplex.dataScanJobs.get on data scan job
    • dataplex.dataScanJobs.list on data scan
  • Publish data profile scan results to Knowledge Catalog:
    • dataplex.entryGroups.useDataProfileAspect on entry group
    • bigquery.tables.update on table
    • dataplex.entries.update on entry
  • View published data profile results for a table in BigQuery or Knowledge Catalog:
    • bigquery.tables.get on table
    • bigquery.tables.getData on table

You might also be able to get these permissions with custom roles or other predefined roles.

Knowledge Catalog service account roles and permissions

Whichever execution identity you select (the default Knowledge Catalog Service Agent, a custom service account, or End-User Credentials), that identity requires the following roles and permissions to run the data profile scan jobs in the backend and export results.

To ensure that the execution identity has the necessary permissions to run data profile scans and export results, ask your administrator to grant the following IAM roles to the execution identity:

  • Run data profile scans:
  • Run data profile scans for BigQuery external tables that use Cloud Storage data:
  • Run data profile scans for Iceberg REST Catalog, SAP BDC Delta Lake, and Apache Hive tables on Google Cloud Lakehouse: BigLake Viewer (roles/biglake.viewer) on the tables being scanned
  • Export data profile scan results to a BigQuery table: BigQuery Data Editor (roles/bigquery.dataEditor) on the table

For more information about granting roles, see Manage access to projects, folders, and organizations.

These predefined roles contain the permissions required to run data profile scans and export results. To see the exact permissions that are required, expand the Required permissions section:

Required permissions

The following permissions are required to run data profile scans and export results:

  • Run data profile scans against BigQuery data:
    • bigquery.jobs.create on project
    • bigquery.tables.get on table
    • bigquery.tables.getData on table
  • Run data profile scans for BigQuery external tables that use Cloud Storage data:
    • storage.buckets.get on bucket
    • storage.objects.get on object
  • Export data profile scan results to a BigQuery table:
    • bigquery.tables.create on dataset
    • bigquery.tables.updateData on table

Your administrator might also be able to give the execution identity these permissions with custom roles or other predefined roles.

If a table uses BigQuery row-level security, then Knowledge Catalog can only scan rows visible to the Knowledge Catalog service account. To let Knowledge Catalog scan all rows, add its service account to a row filter where the predicate is TRUE.

If a table uses BigQuery column-level security, then Knowledge Catalog requires access to scan protected columns. To grant access, give the Knowledge Catalog service account the Data Catalog Fine-Grained Reader (roles/datacatalog.fineGrainedReader) role on all policy tags used in the table. The user creating or updating a data scan also needs permissions on protected columns.

Grant roles to the Knowledge Catalog service account

To run data profile scans, Knowledge Catalog uses a service account that requires permissions to run BigQuery jobs and read BigQuery table data. To grant the required roles, follow these steps:

  1. Get the Knowledge Catalog service account email address. If you haven't created a data profile or data quality scan in this project before, run the following gcloud command to generate the service identity:

    gcloud beta services identity create --service=dataplex.googleapis.com
    

    The command returns the service account email, which has the following format: service-PROJECT_ID@gcp-sa-dataplex.iam.gserviceaccount.com.

    If the service account already exists, you can find its email by viewing principals with the Dataplex name on the IAM page in the Google Cloud console.

  2. Grant the service account the BigQuery Job User (roles/bigquery.jobUser) role on your project. This role lets the service account run BigQuery jobs for the scan.

    gcloud projects add-iam-policy-binding PROJECT_ID \
        --member="serviceAccount:service-PROJECT_NUMBER@gcp-sa-dataplex.iam.gserviceaccount.com" \
        --role="roles/bigquery.jobUser"
    

    Replace the following:

    • PROJECT_ID: your Google Cloud project ID.
    • service-PROJECT_NUMBER@gcp-sa-dataplex.iam.gserviceaccount.com: the email of the Knowledge Catalog service account.
  3. Grant the service account the BigQuery Data Viewer (roles/bigquery.dataViewer) role for each table that you want to profile. This role grants read-only access to the tables.

    gcloud bigquery tables add-iam-policy-binding DATASET_ID.TABLE_ID \
        --member="serviceAccount:service-PROJECT_NUMBER@gcp-sa-dataplex.iam.gserviceaccount.com" \
        --role="roles/bigquery.dataViewer"
    

    Replace the following:

    • DATASET_ID: the ID of the dataset containing the table.
    • TABLE_ID: the ID of the table to profile.
    • service-PROJECT_NUMBER@gcp-sa-dataplex.iam.gserviceaccount.com: the email of the Knowledge Catalog service account.

Networking requirements

To run a scan, you must enable Private Google Access on the VPC subnet that you use for the scan. If you don't specify a subnet, make sure your default subnet has Private Google Access enabled.

Caller identity and execution identity

When you create and run a data profile scan—including in cross-project setups where the scan is in one project and the target resource is in another project—Knowledge Catalog separates operations between two identities that require different permissions:

  • Caller identity (management plane): The user account or service account that sends the API request to create or run the data scan (dataScans.create or dataScans.run).

    To verify that the target resource exists and to validate its schema, the caller identity needs metadata access permission on the target resource (for example, bigquery.tables.get permission on the target BigQuery table). The caller identity doesn't read table rows and doesn't need data access permissions.

  • Execution identity (data plane): The identity that executes the backend compute job, reads the table rows, and generates the profile statistics. By default, Knowledge Catalog uses the Knowledge Catalog Service Agent as the execution identity.

    To read the data in the target tables and generate summaries, the execution identity needs full data access (for example, bigquery.tables.getData) and compute permissions (for example, bigquery.jobs.insert).

Configure a custom execution identity

By default, data profile scans run using the Knowledge Catalog Service Agent. You can override this to use a custom service account or your own End-User Credentials (EUC).

Using a custom execution identity changes how you are billed for the scan. When you specify a custom execution identity, the compute and storage costs associated with the scan are billed directly to your BigQuery project, bypassing the standard Knowledge Catalog Premium SKUs.

Required permissions for custom execution identities

To configure a custom service account or use end-user credentials, you must have the following additional IAM permissions:

  • To use a custom service account, you need the following permissions:
    • The iam.serviceAccounts.actAs permission granted for the project that contains the service account (for example, roles/iam.serviceAccountUser).
    • Your project's Service Agent (service-PROJECT_NUMBER@gcp-sa-dataplex.iam.gserviceaccount.com) needs the iam.serviceAccounts.getAccessToken permission on the custom service account (for example, by having the roles/iam.serviceAccountTokenCreator role).
    • The custom service account needs bigquery.tables.getData on the table to scan, bigquery.jobs.insert in the scan project, and bigquery.dataEditor on the export dataset (if using export).
  • To use End-User Credentials, you need:
    • bigquery.tables.getData on the table to scan.
    • bigquery.jobs.insert in the scan project.
    • bigquery.dataEditor on the export dataset (if using export).

To configure the execution identity, select one of the following options:

Console

To configure the execution identity in the Google Cloud console, select the identity when you create your data profile scan.

In the Execution Identity section, select one of the following:

  • Dataplex service account: The default behavior.
  • Specific service account: Enter the email address of the service account that you want to use.
  • User Credentials: Use your own credentials to run the scan.

REST

To use a custom service account, add the executionIdentity object to your DataScan resource definition during the create request:

"executionIdentity": {
  "serviceAccount": {
     "email": "YOUR_SERVICE_ACCOUNT_EMAIL"
  }
}
  

Replace the following:

  • YOUR_SERVICE_ACCOUNT_EMAIL: the email address of the service account that you want to use.

To use end-user credentials, specify the userCredential object instead:

"executionIdentity": {
  "userCredential": {}
}
  

Create a data profile scan

Console

  1. In the Google Cloud console, go to the Knowledge Catalog Data profiling & quality page.

    Go to Data profiling & quality

  2. Click Create data profile scan.

  3. Optional: Enter a Display name.

  4. Enter an ID. See the Resource naming conventions.

  5. Optional: Enter a Description.

  6. In the Table field, click Browse. Choose the table to scan, and then click Select. Only standard BigQuery, Iceberg REST Catalog, SAP BDC Delta Lake, and Apache Hive on Google Cloud Lakehousetables are supported.

    For tables in multi-region datasets, choose a region where to create the data scan.

    To browse the tables organized within Knowledge Catalog lakes, click Browse within Knowledge Catalog Lakes.

  7. In the Mode section, select one of the following options:

    • Standard: profiles your data with customizable scan settings. This is the default mode.

    • Lightweight: provides quick insights with a low-latency, low-fidelity scan.

  8. If you chose the Standard mode, configure the following options. These options don't appear when you select Lightweight mode.

    1. In the Scope field, choose Incremental or Entire data.

      If you choose Incremental data, in the Timestamp column field, select a column of type DATE or TIMESTAMP from your BigQuery table. Knowledge Catalog uses this column to identify new records as they're added. For tables partitioned on a column of type DATE or TIMESTAMP, it's recommended to use this column as the partition column.

    2. Optional: To filter your data, do any of the following:

      • To filter by rows, select the Filter rows checkbox. Enter a valid SQL expression that can be used in a WHERE clause in GoogleSQL syntax. For example: col1 >= 0.

        The filter can be a combination of SQL conditions over multiple columns. For example: col1 >= 0 AND col2 < 10.

      • To filter by columns, select the Filter columns checkbox.

      • To include columns in the profile scan, in the Include columns field, click Browse. Select the columns to include, and then click Select.

      • To exclude columns from the profile scan, in the Exclude columns field, click Browse. Select the columns to exclude, and then click Select.

    3. To apply sampling to your data profile scan, in the Sampling size list, select a sampling percentage. Choose a percentage value that ranges between 0.0% and 100.0% with up to 3 decimal digits.

      • For larger datasets, choose a lower sampling percentage. For example, for a 1 PB table, if you enter a value between 0.1% and 1.0%, the data profile samples between 1-10 TB of data.

      • There must be at least 100 records in the sampled data to return a result.

      • For incremental data scans, the data profile scan applies sampling to the latest increment.

  9. Optional: Publish the data profile scan results in the BigQuery and Knowledge Catalog pages in the Google Cloud console for the source table. Select the Publish results to Knowledge Catalog checkbox.

    You can view the latest scan results in the Data profile tab in the BigQuery and Knowledge Catalog pages for the source table. To let users access the published scan results, see the Grant access to data profile scan results section of this document.

    The publishing option might not be available in the following cases:

    • You don't have the required permissions on the table.
    • Another data profile scan is set to publish results.
  10. In the Schedule section, choose one of the following options:

    • Repeat: Run the data profile scan on a schedule: hourly, daily, weekly, monthly, or custom. Specify how often the scan should run and at what time. If you choose custom, use cron format to specify the schedule.

    • On-demand: Run the data profile scan on demand.

    • One-time run: Run the data profile scan once now, and remove the scan after the auto-deletion time. This feature's in Preview.

      • Set post-scan results auto-deletion: The auto-deletion time defines the duration a data profile scan remains active after execution. A data profile scan without a specified auto-deletion time is automatically removed after 24 hours. The auto-deletion time can range from 0 seconds (immediate deletion) to 365 days.
  11. Click Continue.

  12. Optional: Export the scan results to a BigQuery standard table. In the Export scan results to BigQuery table section, do the following:

    1. In the Select BigQuery dataset field, click Browse. Select a BigQuery dataset to store the data profile scan results.

    2. In the BigQuery table field, specify the table to store the data profile scan results. If you're using an existing table, make sure that it's compatible with the export table schema. If the specified table doesn't exist, Knowledge Catalog creates it for you.

  13. Optional: Add labels. Labels are key-value pairs that let you group related objects together or with other Google Cloud resources.

  14. To create the scan, click Create.

    If you set the schedule to on-demand, you can also run the scan now by clicking Run scan.

gcloud

To create a data profile scan, use the gcloud dataplex datascans create data-profile command.

If the source data is organized in a Knowledge Catalog lake, include the --data-source-entity flag:

gcloud dataplex datascans create data-profile DATASCAN \
--location=LOCATION \
--data-source-entity=DATA_SOURCE_ENTITY

If the source data isn't organized in a Knowledge Catalog lake, include the --data-source-resource flag:

gcloud dataplex datascans create data-profile DATASCAN \
--location=LOCATION \
--data-source-resource=DATA_SOURCE_RESOURCE

Replace the following variables:

  • DATASCAN: The name of the data profile scan.
  • LOCATION: The Google Cloud region in which to create the data profile scan.
  • DATA_SOURCE_ENTITY: The Knowledge Catalog entity that contains the data for the data profile scan. For example, projects/test-project/locations/test-location/lakes/test-lake/zones/test-zone/entities/test-entity.
  • DATA_SOURCE_RESOURCE: The name of the resource that contains the data for the data profile scan. For example, //bigquery.googleapis.com/projects/test-project/datasets/test-dataset/tables/test-table.

C#

C#

Before trying this sample, follow the C# setup instructions in the Knowledge Catalog quickstart using client libraries. For more information, see the Knowledge Catalog C# API reference documentation.

To authenticate to Knowledge Catalog, set up Application Default Credentials. For more information, see Set up authentication for a local development environment.

using Google.Api.Gax.ResourceNames;
using Google.Cloud.Dataplex.V1;
using Google.LongRunning;

public sealed partial class GeneratedDataScanServiceClientSnippets
{
    /// <summary>Snippet for CreateDataScan</summary>
    /// <remarks>
    /// This snippet has been automatically generated and should be regarded as a code template only.
    /// It will require modifications to work:
    /// - It may require correct/in-range values for request initialization.
    /// - It may require specifying regional endpoints when creating the service client as shown in
    ///   https://cloud.google.com/dotnet/docs/reference/help/client-configuration#endpoint.
    /// </remarks>
    public void CreateDataScanRequestObject()
    {
        // Create client
        DataScanServiceClient dataScanServiceClient = DataScanServiceClient.Create();
        // Initialize request argument(s)
        CreateDataScanRequest request = new CreateDataScanRequest
        {
            ParentAsLocationName = LocationName.FromProjectLocation("[PROJECT]", "[LOCATION]"),
            DataScan = new DataScan(),
            DataScanId = "",
            ValidateOnly = false,
        };
        // Make the request
        Operation<DataScan, OperationMetadata> response = dataScanServiceClient.CreateDataScan(request);

        // Poll until the returned long-running operation is complete
        Operation<DataScan, OperationMetadata> completedResponse = response.PollUntilCompleted();
        // Retrieve the operation result
        DataScan result = completedResponse.Result;

        // Or get the name of the operation
        string operationName = response.Name;
        // This name can be stored, then the long-running operation retrieved later by name
        Operation<DataScan, OperationMetadata> retrievedResponse = dataScanServiceClient.PollOnceCreateDataScan(operationName);
        // Check if the retrieved long-running operation has completed
        if (retrievedResponse.IsCompleted)
        {
            // If it has completed, then access the result
            DataScan retrievedResult = retrievedResponse.Result;
        }
    }
}

Go

Go

Before trying this sample, follow the Go setup instructions in the Knowledge Catalog quickstart using client libraries. For more information, see the Knowledge Catalog Go API reference documentation.

To authenticate to Knowledge Catalog, set up Application Default Credentials. For more information, see Set up authentication for a local development environment.


//go:build examples

package main

import (
	"context"

	dataplex "cloud.google.com/go/dataplex/apiv1"
	dataplexpb "cloud.google.com/go/dataplex/apiv1/dataplexpb"
)

func main() {
	ctx := context.Background()
	// This snippet has been automatically generated and should be regarded as a code template only.
	// It will require modifications to work:
	// - It may require correct/in-range values for request initialization.
	// - It may require specifying regional endpoints when creating the service client as shown in:
	//   https://pkg.go.dev/cloud.google.com/go#hdr-Client_Options
	c, err := dataplex.NewDataScanClient(ctx)
	if err != nil {
		// TODO: Handle error.
	}
	defer c.Close()

	req := &dataplexpb.CreateDataScanRequest{
		// TODO: Fill request struct fields.
		// See https://pkg.go.dev/cloud.google.com/go/dataplex/apiv1/dataplexpb#CreateDataScanRequest.
	}
	op, err := c.CreateDataScan(ctx, req)
	if err != nil {
		// TODO: Handle error.
	}

	resp, err := op.Wait(ctx)
	if err != nil {
		// TODO: Handle error.
	}
	// TODO: Use resp.
	_ = resp
}

Java

Java

Before trying this sample, follow the Java setup instructions in the Knowledge Catalog quickstart using client libraries. For more information, see the Knowledge Catalog Java API reference documentation.

To authenticate to Knowledge Catalog, set up Application Default Credentials. For more information, see Set up authentication for a local development environment.

import com.google.cloud.dataplex.v1.CreateDataScanRequest;
import com.google.cloud.dataplex.v1.DataScan;
import com.google.cloud.dataplex.v1.DataScanServiceClient;
import com.google.cloud.dataplex.v1.LocationName;

public class SyncCreateDataScan {

  public static void main(String[] args) throws Exception {
    syncCreateDataScan();
  }

  public static void syncCreateDataScan() throws Exception {
    // This snippet has been automatically generated and should be regarded as a code template only.
    // It will require modifications to work:
    // - It may require correct/in-range values for request initialization.
    // - It may require specifying regional endpoints when creating the service client as shown in
    // https://cloud.google.com/java/docs/setup#configure_endpoints_for_the_client_library
    try (DataScanServiceClient dataScanServiceClient = DataScanServiceClient.create()) {
      CreateDataScanRequest request =
          CreateDataScanRequest.newBuilder()
              .setParent(LocationName.of("[PROJECT]", "[LOCATION]").toString())
              .setDataScan(DataScan.newBuilder().build())
              .setDataScanId("dataScanId1260787906")
              .setValidateOnly(true)
              .build();
      DataScan response = dataScanServiceClient.createDataScanAsync(request).get();
    }
  }
}

Python

Python

Before trying this sample, follow the Python setup instructions in the Knowledge Catalog quickstart using client libraries. For more information, see the Knowledge Catalog Python API reference documentation.

To authenticate to Knowledge Catalog, set up Application Default Credentials. For more information, see Set up authentication for a local development environment.

# Copyright 2026 Google LLC
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
#      http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

import google.api_core.exceptions
from google.cloud import dataplex_v1


def create_data_profile_scan_global(
    project_id: str,
    dataset_id: str,
    table_id: str,
    location: str,
) -> None:
    """Creates a Dataplex Data Profile Scan using global API endpoint routing.

    Args:
        project_id (str): Google Cloud project ID where the scan is created.
        dataset_id (str): Target BigQuery dataset ID.
        table_id (str): Target BigQuery table ID to scan.
        location (str): Google Cloud region where serverless compute runs.
    """
    client = dataplex_v1.DataScanServiceClient()

    parent = client.common_location_path(project=project_id, location=location)

    bigquery_table = (
        f"//bigquery.googleapis.com/projects/{project_id}"
        f"/datasets/{dataset_id}/tables/{table_id}"
    )

    data_profile_spec = dataplex_v1.DataProfileSpec(sampling_percent=100.0)

    data_scan = dataplex_v1.DataScan(
        display_name="Global Data Profile Scan",
        description="Regional data profile scan generating automated table statistics.",
        data=dataplex_v1.DataSource(resource=bigquery_table),
        data_profile_spec=data_profile_spec,
    )

    request = dataplex_v1.CreateDataScanRequest(
        parent=parent,
        data_scan=data_scan,
    )

    try:
        operation = client.create_data_scan(request=request)
        print(operation.result())

    except google.api_core.exceptions.AlreadyExists:
        print("A scan with this ID already exists.")
    except google.api_core.exceptions.InvalidArgument as e:
        print(f"Your scan configuration is invalid: {e}")
    except google.api_core.exceptions.GoogleAPIError as e:
        print(f"Unexpected exception: {e}")


Ruby

Ruby

Before trying this sample, follow the Ruby setup instructions in the Knowledge Catalog quickstart using client libraries. For more information, see the Knowledge Catalog Ruby API reference documentation.

To authenticate to Knowledge Catalog, set up Application Default Credentials. For more information, see Set up authentication for a local development environment.

require "google/cloud/dataplex/v1"

##
# Snippet for the create_data_scan call in the DataScanService service
#
# This snippet has been automatically generated and should be regarded as a code
# template only. It will require modifications to work:
# - It may require correct/in-range values for request initialization.
# - It may require specifying regional endpoints when creating the service
# client as shown in https://cloud.google.com/ruby/docs/reference.
#
# This is an auto-generated example demonstrating basic usage of
# Google::Cloud::Dataplex::V1::DataScanService::Client#create_data_scan.
#
def create_data_scan
  # Create a client object. The client can be reused for multiple calls.
  client = Google::Cloud::Dataplex::V1::DataScanService::Client.new

  # Create a request. To set request fields, pass in keyword arguments.
  request = Google::Cloud::Dataplex::V1::CreateDataScanRequest.new

  # Call the create_data_scan method.
  result = client.create_data_scan request

  # The returned object is of type Gapic::Operation. You can use it to
  # check the status of an operation, cancel it, or wait for results.
  # Here is how to wait for a response.
  result.wait_until_done! timeout: 60
  if result.response?
    p result.response
  else
    puts "No response received."
  end
end

REST

To create a data profile scan, use the dataScans.create method.

Export table schema

If you want to export the data profile scan results to an existing BigQuery table, make sure that it is compatible with the following table schema:

Column name Column data type Sub field name (if applicable) Sub field data type Mode Example
data_profile_scan struct/record resource_name string nullable //dataplex.googleapis.com/projects/test-project/locations/europe-west2/datascans/test-datascan
project_id string nullable test-project
location string nullable us-central1
data_scan_id string nullable test-datascan
data_source struct/record resource_name string nullable

Entity case: //dataplex.googleapis.com/projects/test-project/locations/europe-west2/lakes/test-lake/zones/test-zone/entities/test-entity

Table case: //bigquery.googleapis.com/projects/test-project/datasets/test-dataset/tables/test-table

dataplex_entity_project_id string nullable test-project
dataplex_entity_project_number integer nullable 123456789012
dataplex_lake_id string nullable

(Valid only if source is entity)

test-lake

dataplex_zone_id string nullable

(Valid only if source is entity)

test-zone

dataplex_entity_id string nullable

(Valid only if source is entity)

test-entity

table_project_id string nullable dataplex-table
table_project_number int64 nullable 345678901234
dataset_id string nullable

(Valid only if source is table)

test-dataset

table_id string nullable

(Valid only if source is table)

test-table

data_profile_job_id string nullable caeba234-cfde-4fca-9e5b-fe02a9812e38
data_profile_job_configuration json trigger string nullable ondemand/schedule
incremental boolean nullable true/false
sampling_percent float nullable

(0-100)

20.0 (indicates 20%)

row_filter string nullable col1 >= 0 AND col2 < 10
column_filter json nullable {"include_fields":["col1","col2"], "exclude_fields":["col3"]}
job_labels json nullable {"key1":value1}
job_start_time timestamp nullable 2023-01-01 00:00:00 UTC
job_end_time timestamp nullable 2023-01-01 00:00:00 UTC
job_rows_scanned integer nullable 7500
column_name string nullable column-1
column_type string nullable string
column_mode string nullable repeated
percent_null float nullable

(0.0-100.0)

20.0 (indicates 20%)

percent_unique float nullable

(0.0-100.0)

92.5

min_string_length integer nullable

(Valid only if column type is string)

10

max_string_length integer nullable

(Valid only if column type is string)

4

average_string_length float nullable

(Valid only if column type is string)

7.2

min_value float nullable (Valid only if column type is numeric - integer/float)
max_value float nullable (Valid only if column type is numeric - integer/float)
average_value float nullable (Valid only if column type is numeric - integer/float)
standard_deviation float nullable (Valid only if column type is numeric - integer/float)
quartile_lower integer nullable (Valid only if column type is numeric - integer/float)
quartile_median integer nullable (Valid only if column type is numeric - integer/float)
quartile_upper integer nullable (Valid only if column type is numeric - integer/float)
top_n struct/record - repeated value string nullable "4009"
count integer nullable 20
percent float nullable 10 (indicates 10%)

Export table setup

When you export to BigQueryExport tables, follow these guidelines:

  • For the field resultsTable, use the format: //bigquery.googleapis.com/projects/{project-id}/datasets/{dataset-id}/tables/{table-id}.
  • Use a BigQuery standard table.
  • If the table doesn't exist when the scan is created or updated, Knowledge Catalog creates the table for you.
  • By default, the table is partitioned on the job_start_time column daily.
  • If you want the table to be partitioned in other configurations or if you don't want the partition, then recreate the table with the required schema and configurations and then provide the pre-created table as the results table.
  • Make sure the results table is in the same location as the source table.
  • If VPC-SC is configured on the project, then the results table must be in the same VPC-SC perimeter as the source table.
  • If the table is modified during the scan execution stage, then the current running job exports to the previous results table and the table change takes effect from the next scan job.
  • Don't modify the table schema. If you need customized columns, create a view upon the table.
  • To reduce costs, set an expiration on the partition based on your use case. For more information, see how to set the partition expiration.

Create multiple data profile scans

You can configure data profile scans for multiple tables in a BigQuery dataset at the same time by using the Google Cloud console.

  1. In the Google Cloud console, go to the Knowledge Catalog Data profiling & quality page.

    Go to Data profiling & quality

  2. Click Create data profile scan.

  3. Select the Multiple data profile scans option.

  4. Enter an ID prefix. Knowledge Catalog automatically generates scan IDs by using the provided prefix and unique suffixes.

  5. Enter a Description for all of the data profile scans.

  6. In the Dataset field, click Browse. Select a dataset to pick tables from. Click Select.

  7. If the dataset is multi-regional, select a Region in which to create the data profile scans.

  8. In the Mode section, choose one of the following options:

    • Standard: profiles your data with customizable scan settings. This is the default mode.

    • Lightweight: provides quick insights with a low-latency, low-fidelity scan. This feature is in Preview.

  9. If you chose the Standard mode, configure the following settings for the scans. These settings don't appear when Lightweight mode is selected.

    1. In the Scope field, choose Incremental or Entire data.

      If you choose Incremental data, you can select only tables that are partitioned on a column of type DATE or TIMESTAMP.

    2. To apply sampling to the data profile scans, in the Sampling size list, select a sampling percentage.

      Choose a percentage value between 0.0% and 100.0% with up to 3 decimal digits.

  10. Optional: Publish the data profile scan results in the BigQuery and Knowledge Catalog pages in the Google Cloud console for the source table. Select the Publish results to Knowledge Catalog checkbox.

    You can view the latest scan results in the Data profile tab in the BigQuery and Knowledge Catalog pages for the source table. To let users access the published scan results, see the Grant access to data profile scan results section of this document.

  11. In the Schedule section, choose one of the following options:

    • Repeat: Run the data profile scans on a schedule: hourly, daily, weekly, monthly, or custom. Specify how often the scans should run and at what time. If you choose custom, use cron format to specify the schedule.

    • On-demand: Run the data profile scans on demand.

      • One-time run: Run the data profile scan once now, and remove the scan after the auto-deletion time. This feature's in Preview.

        • Set post-scan results auto-deletion: The auto-deletion time defines the duration a data profile scan remains active after execution. A data profile scan without a specified auto-deletion time is automatically removed after 24 hours. The auto-deletion time can range from 0 seconds (immediate deletion) to 365 days.
  12. Click Continue.

  13. In the Choose tables field, click Browse. Choose one or more tables to scan, and then click Select.

  14. Click Continue.

  15. Optional: Export the scan results to a BigQuery standard table. In the Export scan results to BigQuery table section, do the following:

    1. In the Select BigQuery dataset field, click Browse. Select a BigQuery dataset to store the data profile scan results.

    2. In the BigQuery table field, specify the table to store the data profile scan results. If you're using an existing table, make sure that it's compatible with the export table schema. If the specified table doesn't exist, Knowledge Catalog creates it for you.

      Knowledge Catalog uses the same results table for all of the data profile scans.

  16. Optional: Add labels. Labels are key-value pairs that let you group related objects together or with other Google Cloud resources.

  17. To create the scans, click Create.

    If you set the schedule to on-demand, you can also run the scans now by clicking Run scan.

Run a data profile scan

Console

  1. In the Google Cloud console, go to the Knowledge Catalog Data profiling & quality page.

Go to Data profiling & quality

  • Click the data profile scan to run.
  • Click Run now.