HOW-TO · BIGQUERY

Catch bad BigQuery data before anyone else.

By Chris Davidson, founder of yForest · Updated September 26, 2026

Write the data quality check as a query, watch it with a Greater than 0 condition, and hear about a problem before a downstream report does.

Start 14-day free trial Download on theMac App Store

No credit card. 14 days. Cancel in one click.

macOS 15+ · Apple Silicon native · 14-day free trial · No credit card

Quick answer: Write a BigQuery query that returns a count representing a data quality problem, duplicate keys, unexpected nulls, values out of range, then click Watch this. Choose A value, set Condition to Greater than 0, pick a Check every interval and destinations, and save. Any BigQuery connection works; no special setup beyond the connection itself.

Before you start

A working BigQuery connection and QueryFlow Pipelines. This is the general Watch This feature applied specifically to data quality checks on BigQuery.

Step by step

  1. Run a data quality check as a query in the SQL Editor against your BigQuery connection.
  2. Click Watch this on the result.
  3. Choose Row count or A value depending on what the check returns.
  4. Set a Condition that represents “bad,” like Greater than 0 for a count of failing rows.
  5. Set Check every, pick Destinations, and Save.

A worked example: catching duplicate keys

On a BigQuery table that should have one row per order ID:

SELECT COUNT(*) AS duplicate_ids
FROM (
  SELECT order_id
  FROM `my-gcp-project.sales.orders`
  WHERE DATE(_PARTITIONTIME) = CURRENT_DATE()
  GROUP BY order_id
  HAVING COUNT(*) > 1
);

Watch A value on duplicate_ids, Condition Greater than 0, Check every 4h. Any duplicate that slips in from a retried load shows up the same day, not weeks later when a report's totals stop matching finance's numbers.

A second worked example: unexpected nulls

For a column that should never be null once an order is marked shipped:

SELECT COUNT(*) AS missing_tracking
FROM `my-gcp-project.sales.orders`
WHERE status = 'shipped' AND tracking_number IS NULL;

Same pattern: A value, Greater than 0, checked on whatever cadence matches how often orders ship.

Check it worked

Run preview to confirm the check currently returns zero (or whatever your clean baseline is) before saving. Use Check now afterward to confirm the watch actually runs and a real check fires correctly.

Troubleshooting

If you seeFix
The check itself errors, unrelated to a bad rowConfirm the BigQuery connection is Connected, per connecting BigQuery.
Results over 25 MB with no LIMITDatabricks results need a LIMIT above that size; BigQuery quality checks should return a small aggregated number, not raw rows, which usually avoids this entirely.
Alerts on data quality issues that were already fixedCheck every runs on the interval you set; a fix between checks won't clear an alert already sent, but the next check will show the corrected count.

Keeping checks cheap

BigQuery bills by bytes scanned, so a quality check that scans an entire history table every run costs more than one worth watching. Scope checks to a partition, like _PARTITIONTIME for today, wherever the table supports it, and the check stays cheap enough to run frequently without a surprising line on the bill.

Databricks version

The same pattern works on Databricks. See catch bad Databricks data first for the equivalent setup with catalog-qualified table names.

QueryFlow Studio $9.99/mo · $99/yr
QueryFlow Pipelines $29.99/mo · $199.99/yr

Frequently asked

Does this need a special BigQuery setup?

No, any working BigQuery connection works, the same one you'd use to query the warehouse normally.

What counts as a data quality check here?

Any query that returns a number representing a problem, a row count of nulls where there shouldn't be any, duplicate keys, rows outside an expected range. If you can write it as a query, you can watch it.

Will this catch schema changes upstream?

Only if you write a check for it specifically, like a query that fails or returns an unexpected count when an expected column is missing. Watch This doesn't detect schema drift automatically.

Does the query need to scan the whole table every check?

Not necessarily. A well-written check scoped to recent data, like today's partition, avoids scanning the whole table on every run, worth keeping in mind for cost on very large tables.

Which tier includes this?

QueryFlow Pipelines, the same tier as Watch This generally.

Catch it before finance does.

14-day free trial, no card. Set up your first BigQuery quality check.

Start 14-day free trial

No credit card. 14 days. Cancel in one click.