Write the data quality check as a query, watch it with a Greater than 0 condition, and hear about a problem before a downstream report does.
No credit card. 14 days. Cancel in one click.
Quick answer: Write a BigQuery query that returns a count representing a data quality problem, duplicate keys, unexpected nulls, values out of range, then click Watch this. Choose A value, set Condition to Greater than 0, pick a Check every interval and destinations, and save. Any BigQuery connection works; no special setup beyond the connection itself.
A working BigQuery connection and QueryFlow Pipelines. This is the general Watch This feature applied specifically to data quality checks on BigQuery.
On a BigQuery table that should have one row per order ID:
SELECT COUNT(*) AS duplicate_ids FROM ( SELECT order_id FROM `my-gcp-project.sales.orders` WHERE DATE(_PARTITIONTIME) = CURRENT_DATE() GROUP BY order_id HAVING COUNT(*) > 1 );
Watch A value on duplicate_ids, Condition Greater than 0, Check every 4h. Any duplicate that slips in from a retried load shows up the same day, not weeks later when a report's totals stop matching finance's numbers.
For a column that should never be null once an order is marked shipped:
SELECT COUNT(*) AS missing_tracking FROM `my-gcp-project.sales.orders` WHERE status = 'shipped' AND tracking_number IS NULL;
Same pattern: A value, Greater than 0, checked on whatever cadence matches how often orders ship.
Run preview to confirm the check currently returns zero (or whatever your clean baseline is) before saving. Use Check now afterward to confirm the watch actually runs and a real check fires correctly.
| If you see | Fix |
|---|---|
| The check itself errors, unrelated to a bad row | Confirm the BigQuery connection is Connected, per connecting BigQuery. |
| Results over 25 MB with no LIMIT | Databricks results need a LIMIT above that size; BigQuery quality checks should return a small aggregated number, not raw rows, which usually avoids this entirely. |
| Alerts on data quality issues that were already fixed | Check every runs on the interval you set; a fix between checks won't clear an alert already sent, but the next check will show the corrected count. |
BigQuery bills by bytes scanned, so a quality check that scans an entire history table every run costs more than one worth watching. Scope checks to a partition, like _PARTITIONTIME for today, wherever the table supports it, and the check stays cheap enough to run frequently without a surprising line on the bill.
The same pattern works on Databricks. See catch bad Databricks data first for the equivalent setup with catalog-qualified table names.
No, any working BigQuery connection works, the same one you'd use to query the warehouse normally.
Any query that returns a number representing a problem, a row count of nulls where there shouldn't be any, duplicate keys, rows outside an expected range. If you can write it as a query, you can watch it.
Only if you write a check for it specifically, like a query that fails or returns an unexpected count when an expected column is missing. Watch This doesn't detect schema drift automatically.
Not necessarily. A well-written check scoped to recent data, like today's partition, avoids scanning the whole table on every run, worth keeping in mind for cost on very large tables.
QueryFlow Pipelines, the same tier as Watch This generally.
14-day free trial, no card. Set up your first BigQuery quality check.
No credit card. 14 days. Cancel in one click.