A warehouse built for SQL against structured tables, and a lakehouse built to sit on top of raw files and run SQL, Python and Spark against them. QueryFlow connects to both, the same as everything else.
No credit card. 14 days. Cancel in one click.
Quick answer: Redshift and Databricks are both connectors in QueryFlow. Redshift connects with IAM authentication over the Redshift Data API, no VPC tunnel or IP allowlist required. Databricks connects with a personal access token or a service principal, plus the server hostname and HTTP path from the SQL warehouse's connection details. Both get the same SQL editor, Ask panel, Scheduler and Watches.
Redshift is the choice for a team that's standardized on AWS and wants a managed, columnar warehouse for structured, already-modeled data. It's a good fit when the data is already relational and the team wants predictable SQL performance without managing servers directly.
Databricks is the choice when raw or semi-structured data needs real transformation before it's query-ready, and when the same platform needs to support both SQL analysts and people writing Spark or Python against the same tables. A Databricks SQL warehouse is the SQL-facing layer on top of that, and it's what QueryFlow connects to.
| Redshift | Databricks | |
|---|---|---|
| Authentication | IAM keys, no VPC tunnel or IP allowlist | Personal access token or OAuth service principal |
| Connects via | Redshift Data API | SQL Warehouse HTTP Path |
| Structure | Database → schema → table | Catalog → schema → table (Unity Catalog) |
| Best fit for | Structured, already-modeled relational data | Raw or semi-structured data needing transformation |
| Query engine | Columnar SQL warehouse | Spark SQL warehouse |
| As a scheduled job source | Yes, including with the app closed | Yes, including with the app closed |
| As a Data Sync target | Not in this release | Insert, Update or Upsert via MERGE |
| QueryFlow tier | Studio and Pipelines | Studio and Pipelines |
Databricks works as a Data Sync target with Insert, Update and Upsert in 1.7. Redshift does not, in this release, it's a query source and a scheduled-job source, and results can be delivered to S3, SFTP, a local file or email, but Data Sync itself isn't wired up to write into Redshift yet. If your workflow specifically needs field-mapped writes into a target table, that currently points you at Databricks or BigQuery, not Redshift.
Both speak SQL close enough to standard that the same logical query barely changes:
-- Redshift SELECT region, SUM(revenue) AS total FROM analytics.public.orders WHERE order_date >= DATEADD(day, -7, GETDATE()) GROUP BY region; -- Databricks SELECT region, SUM(revenue) AS total FROM gold.sales.orders WHERE order_date >= current_date() - INTERVAL 7 DAYS GROUP BY region;
The real difference between the two shows up further upstream, in how the orders table got built in the first place, not in a query like this one against an already-clean table. If that table started life as a pile of raw JSON events, Databricks is more likely where the transformation happened; if it was structured from the start, Redshift handling it directly is just as reasonable.
Redshift's Explorer lists databases, then schemas, then tables, the standard relational three levels. Databricks adds Unity Catalog on top: catalog, then schema, then table, which is one level deeper if your Databricks setup uses more than one catalog to separate raw, cleaned and production-ready data. Either way, the right-click actions on a table, an auto-generated SELECT with a LIMIT, Insert table name, or Copy name, work the same, so switching between the two mid-session doesn't mean relearning the sidebar.
Redshift's IAM setup tends to be a one-time lift handled by whoever manages AWS access, after which the connection itself is just picking the right keys. Databricks's setup is arguably simpler for an individual: generate a personal access token yourself, copy two fields off the warehouse's own connection page, and you're connected, no AWS console involved. A service principal for a shared or automated Databricks connection is a bit more setup than a personal token, but still doesn't touch AWS IAM at all.
Pick Redshift when the data is already structured and the team wants a straightforward SQL warehouse on AWS. Pick Databricks when raw data needs real transformation before analysts can use it, or when Python and Spark work needs to sit next to the SQL. Plenty of AWS-based teams run both: Redshift for the modeled marts, Databricks upstream for the messier processing.
Yes. The background helper runs Snowflake, Redshift (with IAM keys), BigQuery and Databricks jobs delivering to S3, SFTP, a local file or email, even with QueryFlow closed.
Not in this release. BigQuery and Databricks are Data Sync targets with Insert, Update and Upsert; Redshift currently works as a source, not a target, for Data Sync.
No. Set a Catalog to scope the connection, or leave it blank to browse whatever the credential can see by default.
It avoids a long-lived password stored anywhere, using short-lived credentials instead, which most security teams prefer. QueryFlow supports it specifically to avoid VPC tunneling and IP allowlisting for the connection.
14-day free trial, no card. Connect whichever your stack already runs.
No credit card. 14 days. Cancel in one click.