You could open a notebook and write a load command. Or you could map the columns and click Run.
No credit card. 14 days. Cancel in one click.
Quick answer: Add the CSV as a connection, then build a Data Sync into a Databricks table under Unity Catalog: map columns, choose Insert, Update or Upsert, and run it, or save it as a recurring job. Pipelines tier.
CSV-to-BigQuery and CSV-to-Snowflake are the common cases, but plenty of teams run their warehouse on Databricks instead, especially once Unity Catalog is the place everyone's dashboards already point at. A vendor export, a one-time backfill, a list from another department, all need the same thing: land in a real Databricks table without a notebook and a manual write command in the way.
A returned-items CSV from a warehouse system needs to land in a Unity Catalog table, updating rows for RMAs that already exist and inserting new ones:
-- target table CREATE TABLE IF NOT EXISTS main.ops.returns ( rma_id STRING, sku STRING, quantity INT, reason STRING, processed_at TIMESTAMP );
Map the CSV's RMA, SKU, Qty, Reason and Date columns to rma_id, sku, quantity, reason and processed_at. Set MODE to Upsert and MATCH ON to rma_id, since an RMA number is unique and the same file sometimes gets sent twice with a correction on the second pass.
Dry Run shows the preview before anything writes. After a real run, compare row counts:
SELECT COUNT(*) AS row_count FROM main.ops.returns;
against the CSV's row count minus the header, and check the job history for anything skipped, usually a quantity column that had a stray comma or currency symbol in a source file that wasn't purely numeric.
| If you see | Fix |
|---|---|
| Table or schema not found | Confirm the catalog and schema exist in Unity Catalog before pointing the sync at them; QueryFlow can create the table but not the schema. |
| Choose at least one key field | Pick a MATCH ON column for Update or Upsert modes. |
| Job fails on write | Check the Databricks connection has write access on that schema, not just read access for querying. |
This is a load, not a merge across many files at once. If the returns system exports a new CSV every day, you'll re-point the same sync at each new file, or better, save it as a job pointed at a fixed path. It also won't reconcile against Databricks results over 25 MB automatically; that limit applies to query results you're reading back out, not to a sync writing rows in.
If this file arrives on a predictable schedule, save the sync as a job rather than rebuilding the mapping from scratch each time. The mapping and MODE stay fixed; only the file changes underneath. A structural change to the source, a renamed or reordered column, is the one thing that should send you back to rebuild it rather than trust the old mapping blindly.
A CSV has no native types, everything arrives as text, so the target Databricks column decides how it gets interpreted. A quantity column mapped to an INT target will reject a value with a stray comma or currency symbol rather than silently truncating it, which is usually what you want, an obvious failure beats a quietly wrong number. Date columns are the other common snag: if the source file mixes MM/DD/YYYY and YYYY-MM-DD formats across rows, which happens more often than it should when a file gets hand-edited, normalize the format before syncing rather than hoping the target column parses both.
A one-time backfill and a recurring feed from the same source system usually want different MODEs even though the mapping is identical. The initial historical load is Insert, since nothing exists in the target table yet and there's nothing to match against. Once that's done and the same warehouse system starts sending a new file weekly, switch the saved sync to Upsert with a MATCH ON key, so the ongoing feed corrects and extends the table instead of duplicating everything the backfill already loaded.
See also: Integrations Sync BigQuery to Databricks Sync MySQL to Databricks..
No, you can create it from the sync builder as long as the catalog and schema already exist in Unity Catalog.
Three-part Unity Catalog naming: catalog.schema.table, for example main.ops.returns.
Yes, the 25 MB result limit applies to reading query results back out, not to a sync writing rows into a table.
Rebuild the mapping. A renamed or reordered source column won't automatically match the old mapping.
14-day free trial, no card. Map the CSV once, run it whenever the file shows up.
No credit card. 14 days. Cancel in one click.