HOW-TO · DATA SYNC

Move MySQL rows into Databricks on a schedule.

By Chris Davidson, founder of yForest · Updated September 26, 2026

Operational MySQL data feeding a Databricks notebook or ML pipeline usually means someone exporting a CSV by hand, or a bespoke Spark job nobody wants to own. Data Sync maps the fields once and runs on a schedule instead.

Start 14-day free trial Download on theMac App Store

No credit card. 14 days. Cancel in one click.

macOS 15+ · Apple Silicon native · 14-day free trial · No credit card

Quick answer: Add a MySQL connection and a Databricks connection, open Pipelines, Build, New Sync, and pick MySQL as source, Databricks as target. Map fields to a catalog.schema.table path, choose Insert, Update or Upsert, Dry Run, then save as a job. Pipelines tier.

Why MySQL data ends up in Databricks

Application data lives in MySQL. A model that predicts churn or scores a lead needs that data joined against other sources inside a notebook, or a feature engineering job needs it at a scale a MySQL replica isn't built to serve well. The data itself hasn't changed, only where the work happens with it, and getting it there without a hand-rolled export is the actual problem.

Before you start

Steps

  1. Add both connections under Databases.
  2. Open Pipelines → Build and click New Sync.
  3. On the left, pick the MySQL connection and table, or write SQL.
  4. On the right, pick the Databricks connection and target table (catalog.schema.table).
  5. Drag fields across, or click AI Map.
  6. Pick a MODE: Insert, Update or Upsert.
  7. For Update or Upsert, choose a MATCH ON field that's unique per row.
  8. Click Dry Run, then Save as a job.

A worked example

Feeding a churn model with account activity from a MySQL production replica:

SELECT account_id, last_login_at, plan, seats_used, seats_purchased
FROM app.accounts
WHERE last_login_at >= NOW() - INTERVAL 1 DAY OR updated_at >= NOW() - INTERVAL 1 DAY;

Target main.ml_features.account_activity, MODE Upsert, MATCH ON account_id. QueryFlow handles the merge into Databricks on that key, so a login or a seat count change updates the existing row rather than adding a duplicate. Schedule it hourly if the model retrains daily and needs same-day activity.

A type gotcha worth knowing

MySQL commonly represents a boolean as TINYINT(1), an integer that happens to only ever hold 0 or 1, while Databricks has a real BOOLEAN type. If a field like is_active lands as an integer instead of true or false, check how the target column is typed and adjust the mapping. This is a MySQL quirk, not a Data Sync bug, and it shows up in any MySQL-to-warehouse sync, not just this one.

What this doesn't do

This is a scheduled batch sync, not change data capture off the MySQL binlog. If a downstream model genuinely needs sub-minute freshness on every write, a CDC tool reading the binlog directly is the right answer, not a scheduled query re-run every few minutes. For a model that retrains hourly or daily, which covers most real ML workflows, the scheduled approach is simpler to reason about and doesn't require binlog access on the MySQL side at all.

Picking a sync interval that matches the model's retrain cadence

There's rarely a reason to sync more frequently than the model consuming the data actually retrains. A model that retrains nightly gains nothing from an hourly sync except more MERGE operations against a table nobody reads between runs. Match the two schedules and the sync stops being a resource question at all.

Keeping the feature table's schema stable

Data Sync maps exactly the columns you tell it to map. If a data scientist adds a new feature column to the MySQL side, say a computed engagement score, it won't appear in the Databricks table automatically, the mapping needs updating first. Treat a new feature as a reason to revisit the sync, not something that flows through on its own.

Why not just query MySQL from Databricks directly

Databricks can query MySQL directly through a JDBC connection for one-off exploration, and that's often the right move for a single ad hoc question. It's a worse fit for anything a notebook or job runs repeatedly, since every run puts read load on production MySQL and ties the notebook's reliability to MySQL being reachable at exactly that moment. A materialized, synced copy in Unity Catalog decouples the two.

A production pattern worth copying

Teams running this well tend to land the raw synced table first, account_activity as above, untouched, then build the actual feature engineering as a separate notebook or SQL transform reading from it. That keeps the sync itself simple and debuggable, one job, one mapping, while the harder feature logic lives in Databricks where it belongs and can change without touching the sync at all.

A note on the MySQL replica, not the primary

Pointing Data Sync at a MySQL read replica rather than the primary database is worth doing deliberately if the account has one available. An hourly sync reading a modest amount of data won't meaningfully load a well-sized primary, but a replica removes the question entirely, and it's the same practice any analytics workload against production MySQL should already follow.

Versioning the feature definition alongside the sync

A feature table's definition, which columns, which time window, which filters, tends to drift over time as the model owner refines it, and a sync job's mapping quietly encodes those decisions without necessarily documenting why they were made. Keeping the source query and mapping choices noted somewhere the model owner can reference, even a plain comment in the query itself, saves a future debugging session where the model's behavior changed and nobody can tell whether the data or the model logic moved.

What happens when the model gets retired

A feature table built for one specific model has a way of outliving that model's active use, quietly synced hour after hour into a Unity Catalog table nobody queries anymore. Treating this sync the same as any other piece of infrastructure, with an owner and a periodic check that it's still needed, avoids the slow accumulation of syncs that cost nothing to notice and everything to eventually untangle.

What this costs against Fivetran

Fivetran's pricing page (September 26, 2026) meters by monthly active rows: 500,000 MAR free on standard connections, then a $5 base charge per connection between 1 and 1,000,000 MAR, with usage-based pricing above that requiring a quote. QueryFlow Pipelines is $199.99 a year flat, so a feature table refreshed hourly against a few hundred thousand accounts doesn't change the bill at all.

Comparison

QueryFlow Data SyncFivetran
Pricing model$199.99/yr flat (Pipelines)Free under 500,000 MAR, then $5 base + usage quote
Sync patternScheduled batch, Insert/Update/UpsertManaged batch or binlog-based CDC depending on plan
Where it runsYour Mac, on your scheduleManaged cloud service
Setup for this pairTwo connections, one mapped syncConnector wizard, similar order of effort
Best fitOne feature table, small-to-medium volumeMany sources, continuous sync at platform scale

Check it worked

Run Dry Run and compare the preview against MySQL. After a real run, check the job's history for row counts and any listed errors, usually a type mismatch.

Troubleshooting

If you seeFix
Choose at least one key fieldPick a MATCH ON column for Update or Upsert.
Duplicate key values in sourceConfirm the MySQL key is actually unique across the rows being synced.
Boolean field looks numeric in DatabricksMySQL TINYINT(1) case above; adjust the target column type.

Sources

Sync MySQL to BigQuery Databricks Mac client Every warehouse, one client Integrations Sync BigQuery to Databricks Load a CSV into Databricks.
QueryFlow Studio $9.99/mo · $99/yr
QueryFlow Pipelines $29.99/mo · $199.99/yr

Frequently asked

Does this need a Databricks cluster running, or just a SQL warehouse?

A SQL warehouse is enough for Data Sync writes; you don't need an all-purpose cluster unless something else in your workflow, like a notebook, needs one.

How does Upsert write into a Unity Catalog table?

It runs a MERGE on the MATCH ON column you set, inserting new rows and updating existing ones in the same operation, the standard approach for both BigQuery and Databricks as Data Sync targets.

A MySQL TINYINT(1) boolean landed as a number in Databricks. Why?

MySQL commonly stores booleans as TINYINT(1), a plain integer under the hood, while Databricks has a real BOOLEAN type. Check the mapped column's type on the Databricks side and adjust if a true/false field looks numeric.

Which tier includes Data Sync?

Pipelines. Studio covers querying MySQL and Databricks, not building a sync between them.

Feed Databricks from MySQL without a script.

14-day free trial, no card. Map it once, run it hourly.

Start 14-day free trial

No credit card. 14 days. Cancel in one click.