HOW-TO · DATA SYNC

Sync Postgres into Databricks.

By Chris Davidson, founder of yForest · Updated September 26, 2026

Operational data lives in Postgres. Analysis happens in Databricks. Data Sync moves rows from one to the other on a schedule, without a custom pipeline.

Start 14-day free trial Download on theMac App Store

No credit card. 14 days. Cancel in one click.

macOS 15+ · Apple Silicon native · 14-day free trial · No credit card

Quick answer: Open Pipelines, Build, and New Sync. Pick Postgres as the source connection and table, Databricks as the target, and drag fields across, or use AI Map. Choose Insert, Update or Upsert, run a Dry Run to preview, then save it as a scheduled job. Pipelines tier.

Why not just query Postgres from Databricks directly

You can federate a live query, but for anything you'll run repeatedly, a dashboard, a model input, a report someone checks daily, copying the data into Databricks on a schedule is usually simpler and faster than querying across systems every time. Data Sync handles that copy: mapped once, repeatable, with a mode that decides what happens to rows that already exist.

Before you start

Steps

  1. Open Pipelines → Build and click New Sync.
  2. On the left, pick the source Connection (Postgres) and Table, or write SQL.
  3. On the right, pick the target Connection (Databricks) and Table.
  4. Drag from a source field to a target field, or click AI Map to let it guess the pairing.
  5. Pick a MODE: Insert, Update or Upsert.
  6. For Update or Upsert, choose a MATCH ON field whose values are unique per row.
  7. Click Dry Run to preview without writing anything.
  8. Save it as a job, or run it now.

A worked example

Moving a customers table from Postgres into a Databricks silver layer, keeping it current as records change:

SELECT customer_id, email, plan, updated_at
FROM public.customers
WHERE updated_at >= now() - interval '1 day';

Map customer_id to the target's customer_id, set MODE to Upsert, and choose customer_id as MATCH ON. QueryFlow handles the MERGE into Databricks on your key column, no hand-written MERGE statement needed. Save it as a job running every hour to keep the silver table current without a full reload each time.

Keeping schemas aligned on both sides

Data Sync maps the fields you tell it to map; it won't add a new Postgres column to the Databricks target automatically. If the source table gains a column you want mirrored, add it to the Databricks table first and update the mapping. Treat schema changes as a manual step on both ends.

Handling a larger table the first time

For an initial sync of a big table, run it once manually with Dry Run, then a real run, before turning it into a recurring job. Watching the first full run land correctly is worth the extra few minutes before a schedule starts running it unattended.

Check it worked

Run Dry Run first and read the preview before anything writes. After a real run, check the job's history: a clean run shows rows synced with no errors; a partial one lists which rows failed, usually a type mismatch.

Troubleshooting

If you seeFix
Choose at least one key fieldPick a MATCH ON column for Update or Upsert.
Duplicate key values in sourcePick MATCH ON columns that uniquely identify rows in Postgres.
Records synced, with errors listedUsually a type mismatch; read the listed error for the exact field.

For the general Data Sync walkthrough, see the Sync data into BigQuery or Databricks tutorial.

Naming the sync so it's findable later

A sync named "Sync 2" is hard to recognize months later. Name it for what it actually moves, "Postgres customers to Databricks silver," so anyone looking at the Pipelines list, including future you, knows what it does without opening it.

Handling a schema that drifts over time

If the Postgres source gains new columns you'd like mirrored into Databricks, the mapping doesn't pick them up automatically, add the field on the Databricks side and update the mapping by hand. The sync is a mapping you set deliberately, not something that tracks schema drift on its own.

Testing a mapping change before it runs live

After updating a mapping to include a new field, run Dry Run again before the next scheduled run goes out unattended. It costs a minute and catches a mismatched type before it becomes a job failure at 3 AM instead of a preview you caught in daylight.

Where Upsert earns its keep here

Postgres to Databricks syncs are a common case for Upsert specifically, operational records change in place in Postgres, and Insert alone would just pile up duplicates in Databricks every run. Match on the primary key you already trust in Postgres and the MERGE handles the rest.

Related syncs

See also: Integrations Sync BigQuery to Databricks Load a CSV into Databricks..

Integrations Sync BigQuery to Databricks Load a CSV into Databricks.
QueryFlow Studio $9.99/mo · $99/yr
QueryFlow Pipelines $29.99/mo · $199.99/yr

Frequently asked

Does this run once or on a schedule?

Either. Run it manually from Build, or save it as a job with its own schedule under Pipelines.

How does Upsert actually write into Databricks?

It uses a MERGE on the key column you set as MATCH ON, inserting new rows and updating existing ones in a single operation.

What happens to rows deleted in Postgres?

Data Sync inserts and updates; it doesn't delete rows from the Databricks target that disappeared from the source.

Which QueryFlow tier includes Data Sync?

Pipelines. Studio covers querying and the explorer; Data Sync, scheduling and Watch This are Pipelines features.

Move it once, keep it moving.

14-day free trial, no card. Both connections, one mapped sync.

Start 14-day free trial

No credit card. 14 days. Cancel in one click.