By Chris Davidson, founder of yForest · Updated September 26, 2026
Operational MySQL data feeding a Databricks notebook or ML pipeline usually means someone exporting a CSV by hand, or a bespoke Spark job nobody wants to own. Data Sync maps the fields once and runs on a schedule instead.
No credit card. 14 days. Cancel in one click.
Quick answer: Add a MySQL connection and a Databricks connection, open Pipelines, Build, New Sync, and pick MySQL as source, Databricks as target. Map fields to a catalog.schema.table path, choose Insert, Update or Upsert, Dry Run, then save as a job. Pipelines tier.
Application data lives in MySQL. A model that predicts churn or scores a lead needs that data joined against other sources inside a notebook, or a feature engineering job needs it at a scale a MySQL replica isn't built to serve well. The data itself hasn't changed, only where the work happens with it, and getting it there without a hand-rolled export is the actual problem.
catalog.schema.table).Feeding a churn model with account activity from a MySQL production replica:
SELECT account_id, last_login_at, plan, seats_used, seats_purchased FROM app.accounts WHERE last_login_at >= NOW() - INTERVAL 1 DAY OR updated_at >= NOW() - INTERVAL 1 DAY;
Target main.ml_features.account_activity, MODE Upsert, MATCH ON account_id. QueryFlow handles the merge into Databricks on that key, so a login or a seat count change updates the existing row rather than adding a duplicate. Schedule it hourly if the model retrains daily and needs same-day activity.
MySQL commonly represents a boolean as TINYINT(1), an integer that happens to only ever hold 0 or 1, while Databricks has a real BOOLEAN type. If a field like is_active lands as an integer instead of true or false, check how the target column is typed and adjust the mapping. This is a MySQL quirk, not a Data Sync bug, and it shows up in any MySQL-to-warehouse sync, not just this one.
This is a scheduled batch sync, not change data capture off the MySQL binlog. If a downstream model genuinely needs sub-minute freshness on every write, a CDC tool reading the binlog directly is the right answer, not a scheduled query re-run every few minutes. For a model that retrains hourly or daily, which covers most real ML workflows, the scheduled approach is simpler to reason about and doesn't require binlog access on the MySQL side at all.
There's rarely a reason to sync more frequently than the model consuming the data actually retrains. A model that retrains nightly gains nothing from an hourly sync except more MERGE operations against a table nobody reads between runs. Match the two schedules and the sync stops being a resource question at all.
Data Sync maps exactly the columns you tell it to map. If a data scientist adds a new feature column to the MySQL side, say a computed engagement score, it won't appear in the Databricks table automatically, the mapping needs updating first. Treat a new feature as a reason to revisit the sync, not something that flows through on its own.
Databricks can query MySQL directly through a JDBC connection for one-off exploration, and that's often the right move for a single ad hoc question. It's a worse fit for anything a notebook or job runs repeatedly, since every run puts read load on production MySQL and ties the notebook's reliability to MySQL being reachable at exactly that moment. A materialized, synced copy in Unity Catalog decouples the two.
Teams running this well tend to land the raw synced table first, account_activity as above, untouched, then build the actual feature engineering as a separate notebook or SQL transform reading from it. That keeps the sync itself simple and debuggable, one job, one mapping, while the harder feature logic lives in Databricks where it belongs and can change without touching the sync at all.
Pointing Data Sync at a MySQL read replica rather than the primary database is worth doing deliberately if the account has one available. An hourly sync reading a modest amount of data won't meaningfully load a well-sized primary, but a replica removes the question entirely, and it's the same practice any analytics workload against production MySQL should already follow.
A feature table's definition, which columns, which time window, which filters, tends to drift over time as the model owner refines it, and a sync job's mapping quietly encodes those decisions without necessarily documenting why they were made. Keeping the source query and mapping choices noted somewhere the model owner can reference, even a plain comment in the query itself, saves a future debugging session where the model's behavior changed and nobody can tell whether the data or the model logic moved.
A feature table built for one specific model has a way of outliving that model's active use, quietly synced hour after hour into a Unity Catalog table nobody queries anymore. Treating this sync the same as any other piece of infrastructure, with an owner and a periodic check that it's still needed, avoids the slow accumulation of syncs that cost nothing to notice and everything to eventually untangle.
Fivetran's pricing page (September 26, 2026) meters by monthly active rows: 500,000 MAR free on standard connections, then a $5 base charge per connection between 1 and 1,000,000 MAR, with usage-based pricing above that requiring a quote. QueryFlow Pipelines is $199.99 a year flat, so a feature table refreshed hourly against a few hundred thousand accounts doesn't change the bill at all.
| QueryFlow Data Sync | Fivetran | |
|---|---|---|
| Pricing model | $199.99/yr flat (Pipelines) | Free under 500,000 MAR, then $5 base + usage quote |
| Sync pattern | Scheduled batch, Insert/Update/Upsert | Managed batch or binlog-based CDC depending on plan |
| Where it runs | Your Mac, on your schedule | Managed cloud service |
| Setup for this pair | Two connections, one mapped sync | Connector wizard, similar order of effort |
| Best fit | One feature table, small-to-medium volume | Many sources, continuous sync at platform scale |
Run Dry Run and compare the preview against MySQL. After a real run, check the job's history for row counts and any listed errors, usually a type mismatch.
| If you see | Fix |
|---|---|
| Choose at least one key field | Pick a MATCH ON column for Update or Upsert. |
| Duplicate key values in source | Confirm the MySQL key is actually unique across the rows being synced. |
| Boolean field looks numeric in Databricks | MySQL TINYINT(1) case above; adjust the target column type. |
A SQL warehouse is enough for Data Sync writes; you don't need an all-purpose cluster unless something else in your workflow, like a notebook, needs one.
It runs a MERGE on the MATCH ON column you set, inserting new rows and updating existing ones in the same operation, the standard approach for both BigQuery and Databricks as Data Sync targets.
MySQL commonly stores booleans as TINYINT(1), a plain integer under the hood, while Databricks has a real BOOLEAN type. Check the mapped column's type on the Databricks side and adjust if a true/false field looks numeric.
Pipelines. Studio covers querying MySQL and Databricks, not building a sync between them.
14-day free trial, no card. Map it once, run it hourly.
No credit card. 14 days. Cancel in one click.