HOW-TO · DATA SYNC

Get a CSV file into a Redshift table.

By Chris Davidson, founder of yForest · Updated September 26, 2026

The standard route for a one-off CSV load into Redshift is the COPY command against a staged S3 file. QueryFlow skips the staging step: map the file's columns and run it.

Start 14-day free trial Download on theMac App Store

No credit card. 14 days. Cancel in one click.

macOS 15+ · Apple Silicon native · 14-day free trial · No credit card

Quick answer: Add the CSV as a file connection, then build a Data Sync: pick the CSV as source and your Redshift table as target. Map columns, choose Insert, Update or Upsert, and run it, or save it as a recurring job. Pipelines tier.

When the COPY command is more setup than the job needs

Redshift's own documentation steers everyone toward COPY for good reason, it's fast and built for volume. But COPY wants the file staged in S3 first, which means an upload step, a bucket and IAM policy that may not already exist for a one-off load, and a COPY statement with the right column list and format options. For a file in the thousands of rows, not millions, that's a lot of infrastructure for a load that happens once or occasionally.

Before you start

Steps

  1. Add the file as a connection: click the + next to Databases and pick CSV/Excel.
  2. Open Pipelines → Build and click New Sync.
  3. On the left, pick the file connection as your source.
  4. On the right, pick your Redshift connection and the target table, or create one.
  5. Drag columns across, or click AI Map to pair them by name.
  6. Pick a MODE: Insert for a fresh load, Update or Upsert for a refresh.
  7. For Update or Upsert, set a MATCH ON field that's unique per row.
  8. Click Dry Run, check the preview, then Save or run it now.

A worked example

A finance team's monthly actuals export needs to land in a Redshift staging table, replacing prior figures for cost centers that already exist:

-- target table
CREATE TABLE IF NOT EXISTS finance.actuals (
  cost_center VARCHAR(32),
  department VARCHAR(64),
  actual_spend DECIMAL(12,2),
  period DATE
);

Map the CSV's CostCenter, Dept, Spend and Period columns to the four target columns, set MODE to Upsert, and MATCH ON to cost_center plus period. Save it as a job if this file arrives every month, and the same sync re-runs against the new file each time.

Check it worked

Dry Run shows exactly what would write before anything does. After a real run, compare row counts between the CSV and the table, and check the job's history for any listed type mismatches.

Troubleshooting

If you seeFix
Choose at least one key fieldPick a MATCH ON column for Update or Upsert modes.
Duplicate key values in sourceCheck the CSV for repeated cost-center-and-period pairs before syncing.
Job fails on writeConfirm the Redshift connection's IAM identity has INSERT and UPDATE on the target table.

When to use COPY instead

A file in the millions of rows, or a load where every second of throughput matters, is still better served by staging in S3 and running a real COPY command. Redshift's COPY is built to parallelize across the cluster's nodes in a way a row-by-row or batched write through a client connection doesn't match at that scale. For the recurring, moderate-sized files most teams actually deal with, month-end actuals, a vendor export, a one-time backfill, Data Sync is simpler and skips the S3 step entirely.

A type gotcha worth knowing

A CSV has no native type system, every value is text until something interprets it. A column that looks like a plain number in the file, a cost center code with leading zeros like 00452, silently loses those zeros if the target column is numeric. Type sensitive identifier-looking columns as VARCHAR on the Redshift side rather than assuming a number-looking column should be numeric.

Reusing the same sync for next month's file

If next month's file has the same column structure, point the same file connection at the new file and re-run the existing sync without rebuilding the mapping. Only a structural change, a renamed or reordered column, calls for revisiting it.

Deciding where the staging table lives

Landing the CSV data in a dedicated staging schema, finance.actuals rather than mixing it directly into a shared reporting schema, keeps a monthly file's quirks, an occasional bad row, a format change, contained to one place. Downstream views or transforms can read from staging and apply whatever cleanup logic the team trusts, rather than every finance dashboard depending directly on however this month's export happened to be formatted.

A note on cluster or workgroup choice

If the Redshift environment is Serverless, a Data Sync write against it draws from the workgroup's configured base RPU capacity the same as any other query would; a provisioned cluster instead draws from whatever's already running. Neither setup needs anything extra configured for QueryFlow specifically, but it's worth knowing which one you're on if a sync seems slower than expected during a period when other heavy queries are also running.

Delimiter and encoding quirks worth catching early

A file exported from a European finance system sometimes uses a semicolon delimiter and a comma as the decimal separator, the reverse of the US convention QueryFlow's CSV parser assumes by default. Run Dry Run against the very first file from a new vendor before scheduling anything, since a delimiter mismatch usually shows up as one giant unparsed column rather than a subtle error, and it's obvious in the preview once you know to look.

Comparing this to a Redshift Spectrum external table

For a file that's queried occasionally rather than loaded into a native table, Redshift Spectrum can read it directly from S3 without a load step at all, which is worth knowing about as an alternative when the goal is ad hoc querying rather than a table other systems join against regularly. Once the file needs to behave like a normal, indexed Redshift table with fast repeated queries against it, loading it in, as this page covers, is the better fit.

Handling a header row that doesn't start on row one

Some exports include a title line or a generated-on timestamp above the actual header row, which throws off automatic column detection. If a file consistently has this shape, note it as a quirk of that specific vendor's export and adjust the source range when setting up the file connection rather than assuming every future file needs manual cleanup before it can be mapped.

What this costs against Fivetran

Fivetran's file-based connectors, like its Google Cloud Storage or S3 file connectors, meter the same way as any other source under its pricing page (September 26, 2026): 500,000 MAR free, then a $5 base charge per connection between 1 and 1,000,000 MAR, usage above that behind a quote. A monthly finance file with a few thousand rows sits well inside the free tier regardless, so the practical difference here is mostly setup time and the S3 staging step Fivetran's own file connectors also typically expect.

Sources

Load a CSV into Snowflake Redshift Mac client Every warehouse, one client Integrations Sync MySQL to Redshift. Sync PostgreSQL to Redshift
QueryFlow Studio $9.99/mo · $99/yr
QueryFlow Pipelines $29.99/mo · $199.99/yr

Frequently asked

Doesn't Redshift normally want the file staged in S3 first?

The standard COPY command does, since it's built for bulk loads from S3, and it's still the right tool for a very large file. QueryFlow's Data Sync writes rows directly through the connection instead, which skips the staging step for a file that isn't huge.

What write access does the Redshift connection need?

INSERT and UPDATE privileges on the target table, on top of the IAM access key the connection already uses to authenticate.

Is there a practical file size limit?

Practical limits come from your Mac's available memory for very large files and general Redshift load performance; a file in the millions of rows is usually better handled with a real COPY from S3 instead.

Can I schedule this if the file arrives weekly?

Yes, save the sync as a job and set a schedule. If the file's location changes each week, update the file connection before each run.

Skip the S3 staging step.

14-day free trial, no card. Map it once, run it whenever the file shows up.

Start 14-day free trial

No credit card. 14 days. Cancel in one click.