HOW-TO · DATA SYNC

Move Snowflake tables into Databricks.

By Chris Davidson, founder of yForest · Updated September 26, 2026

Teams moving analytics workloads onto a Databricks lakehouse still have production Snowflake tables to bring across. Data Sync moves them on a schedule, mapped and MERGEd, without standing up a pipeline tool for a migration that's mostly a one-time job plus a trailing sync.

Start 14-day free trial Download on theMac App Store

No credit card. 14 days. Cancel in one click.

macOS 15+ · Apple Silicon native · 14-day free trial · No credit card

Quick answer: Add working Snowflake and Databricks connections, open Pipelines, Build, New Sync, and pick Snowflake as the source and Databricks as the target. Map fields to a Unity Catalog table (catalog.schema.table), choose Insert, Update or Upsert, Dry Run it, then save as a job. Pipelines tier.

Why this pair comes up

A lakehouse migration rarely starts with every table moving at once. It starts with one team's workload, usually because Databricks notebooks or a machine learning pipeline need the data and nobody wants a second manual export step. The Snowflake tables stay authoritative for a while, the Databricks copy exists to unblock that one team, and eventually more tables follow the same path.

Before you start

Steps

  1. Add both connections under Databases, if you haven't.
  2. Open Pipelines → Build and click New Sync.
  3. On the left, pick the Snowflake connection and table, or write a query.
  4. On the right, pick the Databricks connection and target table (catalog.schema.table), or create one.
  5. Drag fields across, or click AI Map.
  6. Pick a MODE: Insert, Update or Upsert.
  7. For Update or Upsert, set a MATCH ON column that's unique per row.
  8. Click Dry Run, then Save as a job.

A worked example

Bringing a Snowflake customer accounts table across, keeping it current as records change:

SELECT ACCOUNT_ID, COMPANY_NAME, PLAN_TIER, MRR, UPDATED_AT
FROM ANALYTICS.PUBLIC.ACCOUNTS
WHERE UPDATED_AT >= DATEADD(day, -1, CURRENT_TIMESTAMP());

Target main.finance.accounts in Unity Catalog, MODE Upsert, MATCH ON account_id. QueryFlow handles the merge into the Databricks table on that key, so a plan change or an MRR update lands as an update on the existing row instead of a duplicate. Schedule it hourly during the migration window, then drop to daily once Snowflake is no longer the primary write target.

A type gotcha worth knowing

Snowflake's NUMBER is a precision-scale decimal by default, and if a mapped column lands in Databricks as DOUBLE instead of DECIMAL(p,s), rounding on large aggregates can drift slightly. For a financial figure like MRR, set the Databricks column type explicitly to DECIMAL(18,2) rather than trusting an auto-created column's inferred type.

What this doesn't do

Data Sync moves rows on a schedule; it isn't a live replication feed off Snowflake's own change tracking, and it won't migrate views, stored procedures, or Snowflake-native features like tasks and streams. If the goal is a full platform migration with dependency mapping and object-for-object conversion, that's a bigger project than a scheduled sync can cover on its own, and a services-led migration tool or Databricks' own Lakehouse Federation is worth evaluating for that scope. For moving table data on a schedule while both platforms run in parallel, this covers it.

Picking a sync interval for a migration window

Early in a migration, when both platforms are being used side by side and people are actively comparing numbers between them, an hourly schedule keeps the gap small enough that nobody's staring at a stale Databricks table during a live comparison. Once the team has confidence in the Databricks copy and Snowflake becomes more of a fallback than a source of truth, dropping to a nightly run is usually plenty, and it means fewer MERGE operations running against a table nobody's actively querying for freshness.

Keeping schemas aligned as tables evolve

Data Sync maps the columns you tell it to map. If a column gets added to the Snowflake source mid-migration, say a new churn_risk_score column added by a data science team, it won't show up in Databricks automatically. Treat a schema change on the Snowflake side as a trigger to revisit the mapping, not something the sync will quietly pick up on its own next run.

A migration checklist worth following

Teams that get through this cleanly tend to follow the same rough order: pick one team's tables first, not the whole warehouse. Run the sync in parallel with the existing Snowflake-based reporting for at least a few weeks. Compare row counts and a handful of spot-checked values on both sides before anyone starts trusting the Databricks copy for a real decision. Only then move the team's actual queries and dashboards over, and only then consider whether Snowflake for that specific workload gets turned off.

A note on compute cost beyond the subscription

QueryFlow's subscription doesn't change what either warehouse bills you. A Snowflake MERGE-driving SELECT and a Databricks SQL warehouse both consume their own compute credits during a sync, the same as they would for any other query. For a moderate-sized incremental sync, that cost is usually a rounding error next to a warehouse's other workloads, but it's worth watching if the source query itself is expensive to compute, a large join or aggregation rather than a simple filtered SELECT.

What this costs against Fivetran

Fivetran's own pricing page (September 26, 2026) prices connectors by monthly active rows (MAR): rows inserted or updated, including deletes, counted per calendar month. Its free plan covers up to 500,000 MAR for standard connections. Past that, Fivetran applies a $5 base charge to any standard connection using between 1 and 1,000,000 MAR, with consumption pricing above the free tier requiring a sales quote rather than a published number. QueryFlow Pipelines is $199.99 a year, flat, regardless of how many rows the Snowflake-to-Databricks sync moves in a given month.

Comparison

QueryFlow Data SyncFivetran
Pricing model$199.99/yr flat (Pipelines)Free under 500,000 MAR, then $5 base + usage, quote required above the free tier
Where it runsOn your Mac, on a schedule you setManaged cloud service
AuthSnowflake PAT, Databricks PAT + SQL warehouseOAuth or key-based, vendor-managed
Sync patternScheduled batch: full, incremental, or Upsert on a keyManaged batch and log-based CDC depending on connector
SetupTwo connections, one mapped sync, minutesConnector setup through a guided wizard, similar order of effort
Best fitOne team's tables, a migration window, small-to-medium volumeContinuous sync across many sources at platform scale

Why not Airbyte for this instead

Airbyte covers this pair too, and it's a reasonable choice if the team already runs Airbyte for other sources and wants one more connector in the same place. The tradeoff is the same one Airbyte always carries: you're hosting and maintaining it, whether that's Docker Compose on a server or Airbyte Cloud's own subscription. For a migration that's a handful of tables and a few months of parallel running, standing up infrastructure for it is a heavier commitment than the migration itself usually justifies.

Check it worked

Run Dry Run and compare the preview against a spot check in Snowflake. After a real run, check the job's history for row counts and any listed type errors.

Troubleshooting

If you seeFix
Choose at least one key fieldPick a MATCH ON column for Update or Upsert.
Duplicate key values in sourceConfirm the Snowflake key is actually unique; a composite key needs every column in MATCH ON.
Numeric values look rounded in DatabricksSet the target column to DECIMAL with explicit precision instead of trusting an inferred DOUBLE.

Sources

Sync Databricks to Snowflake Sync BigQuery to Databricks Every warehouse, one client Integrations Load a CSV into Databricks.
QueryFlow Studio $9.99/mo · $99/yr
QueryFlow Pipelines $29.99/mo · $199.99/yr

Frequently asked

Why move off Snowflake instead of just adding Databricks alongside it?

Most teams doing this aren't fully leaving Snowflake on day one. A common pattern is running both during a migration window: Data Sync keeps the Databricks copy current while dashboards and jobs still point at Snowflake, then cut over once the Databricks side is trusted.

Does this handle Snowflake VARIANT columns?

A VARIANT column reads as its JSON text representation. Map it to a Databricks STRING column directly, or parse it into typed columns in Flow Books before it lands if you need it queryable as structured data.

What happens to Snowflake-specific objects like streams or tasks?

Nothing. Data Sync moves table data on the schedule you set; it doesn't replicate Snowflake's CDC objects, and any streams or tasks defined in Snowflake keep running independently of this sync.

Which tier includes this?

Pipelines. Studio covers connecting to and querying both Snowflake and Databricks, not building a scheduled sync between them.

Bring Snowflake tables across on your terms.

14-day free trial, no card. Map it once, run it on a schedule.

Start 14-day free trial

No credit card. 14 days. Cancel in one click.