Write the data quality check as a query against your SQL Warehouse, watch it with a Greater than 0 condition, and hear about a problem early.
No credit card. 14 days. Cancel in one click.
Quick answer: Write a Databricks query that returns a count representing a data quality problem, duplicate keys, out-of-range values, then click Watch this. Choose A value, set Condition to Greater than 0, pick a Check every interval and destinations, and save. Works with or without Unity Catalog; the SQL warehouse just needs to be running.
A working Databricks connection with a running SQL warehouse, and QueryFlow Pipelines.
On a Databricks table that should have exactly one row per customer:
SELECT COUNT(*) AS duplicate_customers FROM ( SELECT customer_id FROM main.crm.customers GROUP BY customer_id HAVING COUNT(*) > 1 );
Watch A value on duplicate_customers, Condition Greater than 0, Check every 1d. A dedup bug in an upstream sync shows up the next check instead of surfacing weeks later as a mismatched customer count in a board deck.
For a metrics table where a percentage column should never exceed 100:
SELECT COUNT(*) AS bad_rows FROM main.analytics.daily_metrics WHERE conversion_rate > 100 OR conversion_rate < 0;
Same pattern, A value, Greater than 0, whatever cadence matches how often the table refreshes.
Run preview confirms the check currently returns a clean baseline before you save. Check now afterward confirms a real check fires and the destination receives it.
| If you see | Fix |
|---|---|
| "SQL warehouse not found or not running" | Start the warehouse, then re-check the connection, per connecting Databricks. |
| Results over 25 MB | Add a LIMIT or reduce columns; a quality check should return an aggregate, not raw rows. |
| Unity Catalog permission errors on the check query | Confirm the connection's credentials have SELECT on the tables the check references. |
A recent migration added a required region column, but some rows written by an older job version still come through null:
SELECT COUNT(*) AS missing_region FROM main.sales.orders WHERE region IS NULL AND order_date >= current_date() - INTERVAL 1 DAYS;
Watching this catches straggler writes from an un-updated job long after the migration itself is done, which a one-time backfill check wouldn't.
The same pattern applies on BigQuery. See catch bad BigQuery data first for the equivalent setup.
No, it works whether or not Unity Catalog is set up, the check is just a query against whatever catalog and schema your connection points at.
Databricks results over 25 MB need a LIMIT or fewer columns. A well-written quality check returning an aggregated count shouldn't hit that limit.
Yes, if your connection has access to more than one catalog, a check can query across them using fully qualified catalog.schema.table references.
Yes, same as any query, if the warehouse is stopped, the check fails the same way a manual query would, and QueryFlow's diagnostics would flag the same issue.
QueryFlow Pipelines.
14-day free trial, no card. Set up your first Databricks quality check.
No credit card. 14 days. Cancel in one click.