The TabFM model

This document describes BigQuery's built-in TabFM tabular regression and classification model.

The built-in TabFM model is an implementation of Google Research's open source TabFM model. The Google Research TabFM model is a foundation model for tabular data that enables zero-shot regression and classification on structured data through in-context learning. Because the TabFM model is pre-trained on hundreds of millions of synthetic datasets generated using structural causal models, it captures complex feature interactions and generalizes well to unseen real-world tables across many domains.

You can use the TabFM model with the AI.PREDICT function to perform regression and classification on structured data in a single forward pass without having to train a model, optimize hyperparameters, or engineer features. The prediction results are comparable to conventional supervised tree-based algorithms such as XGBoost and random forests. If you want more model tuning options than the TabFM model offers, you can train a supervised model such as a boosted tree or random forest model and use it with the ML.PREDICT function instead.

To generate predictions with the TabFM model on tabular data, use the AI.PREDICT function.

To evaluate predicted values from the TabFM model against the actual values, use the AI.EVALUATE function.

To learn more about the Google Research TabFM model, use the following resources:

When you use TabFM through BigQuery, your usage is governed by the Google Cloud Terms of Service and allows for commercial uses. The non-commercial license associated with the publicly downloadable TabFM weights on GitHub and Hugging Face applies only to self-hosted downloads and does not restrict usage within BigQuery.