HOW-TO · DATA SYNC

Load a CSV into BigQuery without a script.

By Chris Davidson, founder of yForest · Updated September 26, 2026

You could write a Python script with the BigQuery client library for a one-time load. Or you could map the columns and click Run.

Start 14-day free trial Download on theMac App Store

No credit card. 14 days. Cancel in one click.

macOS 15+ · Apple Silicon native · 14-day free trial · No credit card

Quick answer: Add the CSV as a connection, then build a Data Sync: pick the CSV as the source and your BigQuery table as the target. Map columns, choose Insert, Update or Upsert, and run it, or save it as a recurring job if the file gets refreshed regularly. Pipelines tier.

When a CSV is the source of truth, for now

A vendor sends you a spreadsheet. A one-off export from another system needs to land in your warehouse. Someone in finance has a file that has to get into BigQuery before Monday's report runs. None of this needs a script, it needs a mapped load you can also repeat if the file shows up again next week.

Before you start

Steps

  1. Add the file as a connection: click the + next to Databases and pick CSV/Excel.
  2. Open Pipelines → Build and click New Sync.
  3. On the left, pick the file connection as your source.
  4. On the right, pick your BigQuery connection and the target table, or create one.
  5. Drag columns across, or click AI Map to pair them by name.
  6. Pick a MODE: Insert for a fresh load, Update or Upsert if rows already exist and you're refreshing them.
  7. For Update or Upsert, set a MATCH ON field, an ID or SKU column that's unique per row.
  8. Click Dry Run, check the preview, then Save or run it now.

A worked example

A vendor's weekly inventory file needs to land in a staging table, replacing stale rows for SKUs that already exist:

-- target table
CREATE TABLE IF NOT EXISTS `my-gcp-project.staging.vendor_inventory` (
  sku STRING,
  quantity INT64,
  updated_at TIMESTAMP
);

Map the CSV's SKU, Qty and LastUpdated columns to sku, quantity and updated_at, set MODE to Upsert, and MATCH ON to sku. Save it as a job if this file arrives every week, and the same sync just re-runs against the new file.

Check it worked

Dry Run shows you exactly what would write before anything does. After a real run, check row counts in BigQuery against the source file, and look at the job's history for any listed errors, usually a type mismatch between a text column and a numeric target.

Troubleshooting

If you seeFix
Choose at least one key fieldPick a MATCH ON column for Update or Upsert modes.
Duplicate key values in sourceCheck the CSV for repeated IDs before syncing; dedupe upstream.
Job fails on writeGive the BigQuery connection BigQuery Data Editor, not just Data Viewer.

When to save it as a recurring job instead

If the same vendor sends a fresh file every week at a predictable path, saving the sync as a job and re-pointing the file connection each time beats rebuilding the mapping from scratch. The field mapping stays the same; only the underlying file changes. If the file's structure itself changes column by column, that's a sign to rebuild the mapping rather than trust the old one.

A quick sanity check after loading

Matching row counts between the source file and the target table is the fastest confirmation a load went cleanly:

SELECT COUNT(*) AS row_count
FROM my_gcp_project.staging.vendor_inventory;

If that number doesn't match the CSV's row count minus its header row, check the job history for skipped rows before trusting the table.

When the file changes shape

If a vendor adds or renames a column in a later file, the existing mapping won't automatically pick it up, it will just ignore the new column or fail on a missing one it expected. Treat a structural change in the source file as a reason to revisit the mapping, not something to assume still works.

Reusing the same sync for a different file

If next week's file has the identical column structure, you can point the same file connection at the new file and re-run the existing sync without rebuilding the mapping. It's only a structural change, a renamed or reordered column, that calls for revisiting it.

A note on very large files

For a file in the hundreds of thousands of rows or more, splitting it before loading can be faster than one giant sync, particularly if you're troubleshooting a single bad row somewhere in the middle. A smaller batch that fails is much quicker to debug than one that fails halfway through a much larger load.

Related syncs

See also: Integrations Sync Databricks to BigQuery Load an Excel File into BigQuery..

Integrations Sync Databricks to BigQuery Load an Excel File into BigQuery.
QueryFlow Studio $9.99/mo · $99/yr
QueryFlow Pipelines $29.99/mo · $199.99/yr
Sync Google Sheets to BigQuery Sync MySQL to BigQuery Load an Excel file into BigQuery Load a CSV into Databricks

Frequently asked

Does the CSV need headers matching the target table exactly?

No, you map each column manually or with AI Map. Names don't need to match, only the mapping does.

Can I schedule this if the file updates weekly?

Yes, save the sync as a job and set a schedule. If the file itself changes location each week, you'll need to update the file connection before each run.

What write role does BigQuery need?

BigQuery Data Editor on the service account or Google login used for the connection, in addition to Job User.

Is there a file size limit?

Practical limits come from BigQuery load limits and your Mac's available memory for large files; very large files may be better split before loading.

Skip the script.

14-day free trial, no card. Map it once, run it whenever the file shows up.

Start 14-day free trial

No credit card. 14 days. Cancel in one click.