HOW-TO

Move Databricks tables into BigQuery.

By Chris Davidson, founder of yForest · Updated September 26, 2026

Move finished output, like scored features, from Databricks into BigQuery without replatforming either workload, on a schedule while QueryFlow is open.

Start 14-day free trial Download on theMac App Store

No credit card. 14 days. Cancel in one click.

macOS 15+ · Apple Silicon native · 14-day free trial · No credit card

Quick answer: Add Databricks and BigQuery connections under Databases, then in Pipelines → Build, start a New Sync with Databricks (Unity Catalog catalog.schema.table naming) as source and BigQuery as target. Map fields, pick Insert, Update, or Upsert, and Dry Run before saving. Pipelines tier; the sync runs on its schedule while QueryFlow is open.

Before you start

A working Databricks connection (personal access token or service principal) and a working BigQuery connection, both added under Databases, and a BigQuery dataset that already exists. Pipelines tier. Data Sync jobs run while QueryFlow is open, so keep the app running through your scheduled sync time.

Steps

  1. Open Pipelines → Build and click New Sync.
  2. On the left, pick your Databricks Connection and Table, or write SQL against it.
  3. On the right, pick your BigQuery Connection and target Table.
  4. Drag from a source field to a target field, or click AI Map.
  5. Pick a MODE: Insert, Update, or Upsert.
  6. For Update or Upsert, choose a MATCH ON field.
  7. Click Dry Run, then Save it as a job or run it now.
QueryFlow Data Sync mapper with a Databricks source and BigQuery target
Databricks on the left, BigQuery on the right, one line per mapped field.

A worked example

A company runs its machine-learning feature pipeline on Databricks but keeps its BI dashboards on BigQuery, and wants a scored-output table copied over nightly. The Databricks source, using Unity Catalog's three-level naming:

SELECT
  customer_id,
  churn_score,
  scored_at
FROM prod.ml_features.churn_scores
WHERE scored_at >= current_date() - INTERVAL 1 DAY;

The BigQuery target, project and dataset backticked:

`my-gcp-project.ml_outputs.churn_scores`

Map customer_id, churn_score, and scored_at, set MODE to Upsert, and MATCH ON to customer_id, so each day's re-scoring updates the same customer's row rather than accumulating a new one every run. QueryFlow implements this as a BigQuery MERGE against the matched key.

Check it worked

Dry Run first, watching the row count against how many distinct customers you expect scored that day. Databricks query results over 25 MB need a LIMIT or fewer columns before the sync can pull them cleanly, worth checking on a wide feature table before scheduling it.

Troubleshooting

If you seeFix
No Databricks connections availableAdd one under Databases first (personal access token or service principal).
Job fails on writeGive the BigQuery service account BigQuery Data Editor on the target dataset.
Duplicate key values in source for key column(s) …Confirm your Databricks query returns one row per customer_id for the day; a join that fans out rows will violate the MATCH ON key.

Why teams run this pairing

Databricks' Spark engine and MLflow tooling suit model training and feature engineering; BigQuery's serverless query model suits ad hoc analysis and dashboards that need to scale without anyone managing compute. Rather than replatforming one workload onto the other's engine, Data Sync moves the specific tables that need to cross the boundary, the finished scores, not the training data or the notebooks that produced them.

Cost math against Fivetran's own pricing

Fivetran connects both Databricks and BigQuery as destinations and sources, metered the same MAR way: distinct rows changed per calendar month, per its own published definition, with a $5 base charge per connection and a free allowance up to 500,000 MAR. A daily-refreshed scoring table touching a defined customer base each day generates MAR roughly proportional to that customer count, which can sit comfortably inside Fivetran's free tier for a modest base or move into paid usage for a larger one, per the rates on Fivetran's own pricing page. QueryFlow Pipelines runs $29.99/month or $199.99/year flat regardless of how many customers get scored each day.

Type mapping between Databricks and BigQuery

Databricks typeBigQuery typeNote
INT / BIGINTINT64Direct mapping.
DECIMAL(p,s)NUMERIC or BIGNUMERICUse BIGNUMERIC for precision beyond BigQuery's standard NUMERIC range.
TIMESTAMPTIMESTAMPDatabricks TIMESTAMP is stored with an implicit session time zone; confirm what zone the cluster or warehouse assumes.
ARRAY<T>ARRAY<T>Both support native array types; confirm the element type maps cleanly.
STRUCTSTRUCT (RECORD)Nested struct fields map to BigQuery's RECORD type; flatten specific fields in the source query if the destination table expects flat columns instead.
STRINGSTRINGDirect mapping.

Struct and array types are the ones most likely to need extra thought: both Databricks and BigQuery support nested and repeated fields natively, so a struct-to-struct mapping can work directly, but if the downstream BigQuery table is meant to feed a BI tool that expects flat columns, flattening specific nested fields in the Databricks source query (rather than trying to reshape a nested BigQuery RECORD after the fact) is usually the more maintainable approach.

Databricks-specific considerations

A Databricks SQL Warehouse that's set to auto-suspend after a period of inactivity needs to actually be running (or configured to auto-start on a query) for a scheduled sync to succeed; a warehouse that's suspended and set to require manual start will cause the sync to fail rather than wait, similar to a paused Redshift cluster. Checking the warehouse's auto-resume setting in Databricks before relying on an unattended overnight schedule avoids a predictable failure mode.

Monitoring the sync over time

Feature and scoring pipelines on the Databricks side tend to evolve, a data scientist adds a new feature column, renames an existing one, or changes a scoring model's output range. None of those changes propagate into the BigQuery target automatically; the field mapping stays exactly as drawn until someone opens the sync and updates it. Building a habit of checking the sync whenever the upstream Databricks notebook or job changes catches this before a stale mapping quietly drops a newly-added column from the BigQuery table.

A note on model versioning

If the churn-scoring model itself gets retrained and its score range or calibration shifts, that's invisible to Data Sync entirely, since the sync just moves whatever values the Databricks query returns; tracking which model version produced a given batch of scores (a model_version column included in the source query, for example) is worth building into the pipeline upstream if score comparability across model versions matters downstream.

A note on cost on both sides

Databricks bills for the SQL Warehouse compute a query consumes independent of anything QueryFlow charges, and BigQuery bills separately for the write; a frequent, high-volume sync is worth checking against both platforms' own usage dashboards periodically, the same way you'd monitor any regularly-run query's underlying compute cost, rather than assuming a scheduled job is free simply because the scheduling itself has no marginal cost in QueryFlow.

A note on comparing this to a dedicated MLOps pipeline

Teams with a mature MLOps setup often already have a dedicated feature-store or model-serving pipeline moving scores out of Databricks on its own schedule, in which case this sync is redundant with infrastructure that already exists; it's best suited to a team that has a working Databricks scoring job but hasn't yet built (or doesn't want to maintain) a separate delivery pipeline just to get those scores into a BigQuery-based BI tool, where a straightforward scheduled Data Sync closes that gap without adding a new piece of MLOps infrastructure to maintain.

One last habit worth building: after any change to the scoring notebook's output columns or either connection's credentials, run Dry Run once before trusting the next scheduled execution, rather than discovering a mismatch only when a downstream dashboard shows a gap.

Sources

Related syncs

See also: Integrations Load a CSV into BigQuery on Mac Load an Excel File into BigQuery..

Integrations Load a CSV into BigQuery on Mac Load an Excel File into BigQuery.
QueryFlow Studio $9.99/mo · $99/yr
QueryFlow Pipelines $29.99/mo · $199.99/yr

Frequently asked

Does this work with Unity Catalog?

Yes, set the catalog when you configure the Databricks connection, or leave it blank to browse the workspace's default, and the source table reference uses the catalog.schema.table naming Unity Catalog expects.

What's the row limit on a Databricks source query?

Up to 100,000 rows per query, and results over 25 MB need a LIMIT or fewer columns before QueryFlow can pull them.

Can this run with QueryFlow fully closed?

No. Data Sync always runs while QueryFlow is open. Run Jobs When App Is Closed only covers scheduled queries delivering to S3, SFTP, a local file, or email, not Data Sync.

Does Upsert overwrite the whole churn_scores table or just matched rows?

Just matched rows get updated and new customer_ids get inserted; existing rows for customers not present in that day's Databricks query are left untouched.

Databricks output, landing in BigQuery.

14-day free trial, no card. Runs unattended once scheduled.

Start 14-day free trial

No credit card. 14 days. Cancel in one click.