CDC gets used loosely for anything that syncs only changed rows. The strict version reads a database's transaction log; the common version runs a filtered query on a schedule. They aren't the same thing.
No credit card. 14 days. Cancel in one click.
Quick answer: Change Data Capture identifies and delivers only changed rows instead of re-reading a whole table. Log-based CDC reads a database's transaction log directly and catches every change as it happens; query-based sync (what most scheduled tools, including QueryFlow's Data Sync, actually run) filters by a timestamp or ID column on an interval. Log-based CDC catches intermediate states a scheduled query can miss; query-based sync is simpler to set up and needs no replication slot.
Change Data Capture is a pattern for identifying and delivering only the rows that changed in a source system, instead of re-reading the whole table every time. There are two common ways to implement it. Log-based CDC reads a database's own transaction log (Postgres logical replication, MySQL's binlog, SQL Server's CDC feature) to see every insert, update, and delete as it happens, with no query load on the source table itself. Query-based CDC, sometimes called incremental sync, instead runs a periodic query filtered by a timestamp or an auto-incrementing column, like WHERE updated_at > last_run_time, and treats whatever comes back as the changes. Both are "CDC" in the loose sense people mean when they say the word; only the first is CDC in the strict, log-based sense vendors sometimes market it as.
Log-based CDC catches every change, including ones that happen and get reverted within the same interval, and it puts no read load on the source table since it's reading the log, not the table. It also requires enabling and maintaining replication infrastructure on the source: logical replication slots on Postgres, binlog access and retention on MySQL, and usually a dedicated user with broader replication-level permissions. Query-based sync is simpler to set up and reason about, but it can miss a change that happens and reverts between two runs, and if the source table lacks a reliable updated_at column or the column isn't actually updated on every write, the incremental filter silently misses rows.
QueryFlow's Data Sync uses query-based, scheduled sync patterns, not log-based CDC. You pick a sync interval (5 minutes up to daily or a custom cron schedule) and QueryFlow re-runs the source query on that cadence, using Insert, Update, or Upsert mode against the destination based on a MATCH ON key you choose. For a table with a reliable updated_at or auto-increment column, this behaves like the incremental-sync half of CDC for most practical purposes: only changed rows are picked up on each run, at whatever interval you've set. What it doesn't do is read a database transaction log, so a change that happens and reverts inside one sync interval, or a source table with no reliable timestamp to filter on, won't be captured the way log-based CDC would catch it. If your workload genuinely needs sub-second replication or a guarantee that no intermediate state is ever missed, that's a real reason to reach for a log-based CDC tool instead of a scheduled sync, QueryFlow's or otherwise.
-- Query-based sync (what QueryFlow's Data Sync runs on a schedule) SELECT id, status, updated_at FROM orders WHERE updated_at >= now() - interval '15 minutes'; -- vs. log-based CDC (reads the WAL/binlog directly, no query against the table) -- e.g. Postgres logical replication slot emitting every row-level change -- as it commits, independent of any scheduled query.
If an order's status flips from "pending" to "cancelled" and back to "pending" within a 15-minute window, the query-based approach above only ever sees the final state at query time; log-based CDC would have emitted all three changes as they happened.
A less common third approach worth naming: database triggers that write changed rows into a separate change-log table as they happen, which a downstream process then reads. This avoids needing replication-slot access like log-based CDC, but it adds write overhead to every transaction on the source table (the trigger itself has to run synchronously) and requires maintaining the trigger logic alongside any schema change to the source table. It's most often seen in older systems built before managed log-based CDC tooling was widely available, or in databases where logical replication isn't an option at all.
For most reporting and warehouse-freshness use cases, query-based incremental sync on a short interval is the pragmatic default: it requires no special source-side setup beyond a reliable timestamp column, and the gap between "true" CDC and a five- or fifteen-minute polling interval rarely matters to a dashboard a person checks a few times a day. Log-based CDC earns its added setup complexity specifically when downstream systems need every intermediate state (an audit trail that must capture a value that changed and reverted within a minute), when the source table is too large or too write-heavy for repeated polling queries to be practical, or when true sub-second freshness is a hard requirement rather than a nice-to-have.
Part of the confusion comes from vendors marketing "CDC" as a checkbox feature regardless of which mechanism they've actually implemented underneath, since the term carries more prestige in a product comparison than "scheduled incremental sync" does, even when the underlying mechanism is the same query-based approach described above.
It's query-based incremental sync on a schedule, not log-based CDC. For most reporting and warehouse-freshness use cases the practical difference is small; for workloads needing every intermediate state captured, it isn't a substitute for log-based CDC.
QueryFlow connects to Postgres over its standard wire protocol for queries; it doesn't consume a logical replication slot directly.
Scheduled jobs support intervals down to a few minutes, or a custom cron expression for finer control.
Closer, in that less time passes between checks, but it's still point-in-time query-based sync; an interval, however short, can still miss a change-and-revert that happens entirely between two runs.
14-day free trial, no card. See what interval-based Data Sync looks like on your own tables.
No credit card. 14 days. Cancel in one click.