By Chris Davidson, founder of yForest · Updated September 26, 2026
The standard route for a one-off CSV load into Redshift is the COPY command against a staged S3 file. QueryFlow skips the staging step: map the file's columns and run it.
No credit card. 14 days. Cancel in one click.
Quick answer: Add the CSV as a file connection, then build a Data Sync: pick the CSV as source and your Redshift table as target. Map columns, choose Insert, Update or Upsert, and run it, or save it as a recurring job. Pipelines tier.
Redshift's own documentation steers everyone toward COPY for good reason, it's fast and built for volume. But COPY wants the file staged in S3 first, which means an upload step, a bucket and IAM policy that may not already exist for a one-off load, and a COPY statement with the right column list and format options. For a file in the thousands of rows, not millions, that's a lot of infrastructure for a load that happens once or occasionally.
A finance team's monthly actuals export needs to land in a Redshift staging table, replacing prior figures for cost centers that already exist:
-- target table CREATE TABLE IF NOT EXISTS finance.actuals ( cost_center VARCHAR(32), department VARCHAR(64), actual_spend DECIMAL(12,2), period DATE );
Map the CSV's CostCenter, Dept, Spend and Period columns to the four target columns, set MODE to Upsert, and MATCH ON to cost_center plus period. Save it as a job if this file arrives every month, and the same sync re-runs against the new file each time.
Dry Run shows exactly what would write before anything does. After a real run, compare row counts between the CSV and the table, and check the job's history for any listed type mismatches.
| If you see | Fix |
|---|---|
| Choose at least one key field | Pick a MATCH ON column for Update or Upsert modes. |
| Duplicate key values in source | Check the CSV for repeated cost-center-and-period pairs before syncing. |
| Job fails on write | Confirm the Redshift connection's IAM identity has INSERT and UPDATE on the target table. |
A file in the millions of rows, or a load where every second of throughput matters, is still better served by staging in S3 and running a real COPY command. Redshift's COPY is built to parallelize across the cluster's nodes in a way a row-by-row or batched write through a client connection doesn't match at that scale. For the recurring, moderate-sized files most teams actually deal with, month-end actuals, a vendor export, a one-time backfill, Data Sync is simpler and skips the S3 step entirely.
A CSV has no native type system, every value is text until something interprets it. A column that looks like a plain number in the file, a cost center code with leading zeros like 00452, silently loses those zeros if the target column is numeric. Type sensitive identifier-looking columns as VARCHAR on the Redshift side rather than assuming a number-looking column should be numeric.
If next month's file has the same column structure, point the same file connection at the new file and re-run the existing sync without rebuilding the mapping. Only a structural change, a renamed or reordered column, calls for revisiting it.
Landing the CSV data in a dedicated staging schema, finance.actuals rather than mixing it directly into a shared reporting schema, keeps a monthly file's quirks, an occasional bad row, a format change, contained to one place. Downstream views or transforms can read from staging and apply whatever cleanup logic the team trusts, rather than every finance dashboard depending directly on however this month's export happened to be formatted.
If the Redshift environment is Serverless, a Data Sync write against it draws from the workgroup's configured base RPU capacity the same as any other query would; a provisioned cluster instead draws from whatever's already running. Neither setup needs anything extra configured for QueryFlow specifically, but it's worth knowing which one you're on if a sync seems slower than expected during a period when other heavy queries are also running.
A file exported from a European finance system sometimes uses a semicolon delimiter and a comma as the decimal separator, the reverse of the US convention QueryFlow's CSV parser assumes by default. Run Dry Run against the very first file from a new vendor before scheduling anything, since a delimiter mismatch usually shows up as one giant unparsed column rather than a subtle error, and it's obvious in the preview once you know to look.
For a file that's queried occasionally rather than loaded into a native table, Redshift Spectrum can read it directly from S3 without a load step at all, which is worth knowing about as an alternative when the goal is ad hoc querying rather than a table other systems join against regularly. Once the file needs to behave like a normal, indexed Redshift table with fast repeated queries against it, loading it in, as this page covers, is the better fit.
Some exports include a title line or a generated-on timestamp above the actual header row, which throws off automatic column detection. If a file consistently has this shape, note it as a quirk of that specific vendor's export and adjust the source range when setting up the file connection rather than assuming every future file needs manual cleanup before it can be mapped.
Fivetran's file-based connectors, like its Google Cloud Storage or S3 file connectors, meter the same way as any other source under its pricing page (September 26, 2026): 500,000 MAR free, then a $5 base charge per connection between 1 and 1,000,000 MAR, usage above that behind a quote. A monthly finance file with a few thousand rows sits well inside the free tier regardless, so the practical difference here is mostly setup time and the S3 staging step Fivetran's own file connectors also typically expect.
The standard COPY command does, since it's built for bulk loads from S3, and it's still the right tool for a very large file. QueryFlow's Data Sync writes rows directly through the connection instead, which skips the staging step for a file that isn't huge.
INSERT and UPDATE privileges on the target table, on top of the IAM access key the connection already uses to authenticate.
Practical limits come from your Mac's available memory for very large files and general Redshift load performance; a file in the millions of rows is usually better handled with a real COPY from S3 instead.
Yes, save the sync as a job and set a schedule. If the file's location changes each week, update the file connection before each run.
14-day free trial, no card. Map it once, run it whenever the file shows up.
No credit card. 14 days. Cancel in one click.