HOW-TO · DATA SYNC

Get a CSV file into a Databricks table.

By Chris Davidson, founder of yForest · Updated September 26, 2026

You could open a notebook and write a load command. Or you could map the columns and click Run.

Start 14-day free trial Download on theMac App Store

No credit card. 14 days. Cancel in one click.

macOS 15+ · Apple Silicon native · 14-day free trial · No credit card

Quick answer: Add the CSV as a connection, then build a Data Sync into a Databricks table under Unity Catalog: map columns, choose Insert, Update or Upsert, and run it, or save it as a recurring job. Pipelines tier.

A one-off file, a table someone else already queries

CSV-to-BigQuery and CSV-to-Snowflake are the common cases, but plenty of teams run their warehouse on Databricks instead, especially once Unity Catalog is the place everyone's dashboards already point at. A vendor export, a one-time backfill, a list from another department, all need the same thing: land in a real Databricks table without a notebook and a manual write command in the way.

Before you start

Steps

  1. Add the file as a connection: click the + next to Databases and pick CSV/Excel.
  2. Open Pipelines → Build and click New Sync.
  3. On the left, pick the CSV connection as your source.
  4. On the right, pick your Databricks connection and the target table, or create one under a Unity Catalog schema.
  5. Drag columns across, or click AI Map to pair them by name.
  6. Pick a MODE: Insert for a fresh load, Update or Upsert if rows already exist and you're refreshing them.
  7. For Update or Upsert, set a MATCH ON field, a column that's unique per row.
  8. Click Dry Run, check the preview, then Save or run it now.

A worked example

A returned-items CSV from a warehouse system needs to land in a Unity Catalog table, updating rows for RMAs that already exist and inserting new ones:

-- target table
CREATE TABLE IF NOT EXISTS main.ops.returns (
  rma_id STRING,
  sku STRING,
  quantity INT,
  reason STRING,
  processed_at TIMESTAMP
);

Map the CSV's RMA, SKU, Qty, Reason and Date columns to rma_id, sku, quantity, reason and processed_at. Set MODE to Upsert and MATCH ON to rma_id, since an RMA number is unique and the same file sometimes gets sent twice with a correction on the second pass.

Check it worked

Dry Run shows the preview before anything writes. After a real run, compare row counts:

SELECT COUNT(*) AS row_count FROM main.ops.returns;

against the CSV's row count minus the header, and check the job history for anything skipped, usually a quantity column that had a stray comma or currency symbol in a source file that wasn't purely numeric.

Troubleshooting

If you seeFix
Table or schema not foundConfirm the catalog and schema exist in Unity Catalog before pointing the sync at them; QueryFlow can create the table but not the schema.
Choose at least one key fieldPick a MATCH ON column for Update or Upsert modes.
Job fails on writeCheck the Databricks connection has write access on that schema, not just read access for querying.

What this doesn't do

This is a load, not a merge across many files at once. If the returns system exports a new CSV every day, you'll re-point the same sync at each new file, or better, save it as a job pointed at a fixed path. It also won't reconcile against Databricks results over 25 MB automatically; that limit applies to query results you're reading back out, not to a sync writing rows in.

Making it recurring

If this file arrives on a predictable schedule, save the sync as a job rather than rebuilding the mapping from scratch each time. The mapping and MODE stay fixed; only the file changes underneath. A structural change to the source, a renamed or reordered column, is the one thing that should send you back to rebuild it rather than trust the old mapping blindly.

Type mapping from CSV to a Delta table

A CSV has no native types, everything arrives as text, so the target Databricks column decides how it gets interpreted. A quantity column mapped to an INT target will reject a value with a stray comma or currency symbol rather than silently truncating it, which is usually what you want, an obvious failure beats a quietly wrong number. Date columns are the other common snag: if the source file mixes MM/DD/YYYY and YYYY-MM-DD formats across rows, which happens more often than it should when a file gets hand-edited, normalize the format before syncing rather than hoping the target column parses both.

First load versus an ongoing feed

A one-time backfill and a recurring feed from the same source system usually want different MODEs even though the mapping is identical. The initial historical load is Insert, since nothing exists in the target table yet and there's nothing to match against. Once that's done and the same warehouse system starts sending a new file weekly, switch the saved sync to Upsert with a MATCH ON key, so the ongoing feed corrects and extends the table instead of duplicating everything the backfill already loaded.

Related syncs

See also: Integrations Sync BigQuery to Databricks Sync MySQL to Databricks..

Integrations Sync BigQuery to Databricks Sync MySQL to Databricks.
QueryFlow Studio $9.99/mo · $99/yr
QueryFlow Pipelines $29.99/mo · $199.99/yr

Frequently asked

Does the target table need to exist first?

No, you can create it from the sync builder as long as the catalog and schema already exist in Unity Catalog.

What naming does Databricks use for the target?

Three-part Unity Catalog naming: catalog.schema.table, for example main.ops.returns.

Can I load a CSV with more rows than fit in one query result?

Yes, the 25 MB result limit applies to reading query results back out, not to a sync writing rows into a table.

What if the CSV changes columns next time?

Rebuild the mapping. A renamed or reordered source column won't automatically match the old mapping.

No notebook needed.

14-day free trial, no card. Map the CSV once, run it whenever the file shows up.

Start 14-day free trial

No credit card. 14 days. Cancel in one click.