"Scheduler" gets used loosely. In ETL, it means one specific thing: something that fires a task at a set time and keeps track of whether it ran.
No credit card. 14 days. Cancel in one click.
Quick answer: A job scheduler is the component that decides when a task runs, cron, interval, daily, weekly, or a full cron expression, and records whether it succeeded. In QueryFlow, a scheduled job pairs a query with a trigger and an output destination, and its run history is the record. It's a narrower job than orchestration: a scheduler decides when; it doesn't manage dependencies between many tasks.
A job scheduler is software that runs a defined task at a defined time, or on a repeating interval, without a person clicking a button each time. The task itself, running a query, moving a file, triggering a sync, isn't the scheduler's job. The scheduler's job is deciding when that task fires and keeping a record of whether it actually happened.
Most schedulers support a handful of trigger shapes: run once manually, run every N minutes or hours, run daily or weekly at a fixed time, or run on a full cron expression for anything more specific. What separates a good scheduler from a bare cron entry is what happens around the trigger: does it retry on failure, does it tell you when something goes wrong, and does it catch up a run that was missed because the machine running it was asleep or offline.
A scheduled job in QueryFlow is three things bound together: a query (or a Data Sync pipeline), a Trigger Type with a time, and an output destination. Click Schedule on a query, pick Daily, Weekly, Interval, or Custom Cron, choose where results go, and QueryFlow's scheduler handles the "when." Its run history is the record of whether each firing succeeded, with the scheduled time and the actual run time tracked separately so a job that caught up late after sleep looks different from one that simply failed.
On Pipelines, jobs also get retry behavior and failure alerts layered on top of that base scheduling, and specific source/destination pairs (Snowflake, Redshift with IAM keys, BigQuery, and Databricks jobs delivering to S3, SFTP, a file, or email) keep running via a background helper even with the app closed. Everything else needs the app open at the trigger time.
A scheduler answers "when does this run." An orchestrator (Airflow, Dagster) answers a bigger question: "in what order do these fifty interdependent tasks run, and what happens if task twelve fails." QueryFlow's scheduler is the former, on purpose. For linear chains where one job should run after another, a Flow Book covers most of that need without introducing DAG configuration.
Say you have a nightly job: pull yesterday's orders from Postgres, then push the aggregated total into a Snowflake table finance reads from. A scheduler handles the "pull at 2 AM" part. It does not, on its own, guarantee the push only happens after the pull finished successfully, or retry the push if the pull's output was empty for some unrelated reason. In QueryFlow, that ordering is handled by a Flow Book running both steps in sequence, with the scheduler firing the Flow Book as a single unit rather than firing two independent jobs and hoping they land in the right order.
In casual conversation, people say "the scheduler" to mean the whole system that keeps their data pipelines running: the triggers, the retries, the alerting, sometimes even the dependency logic between jobs. That's understandable, since from a user's seat it often feels like one thing. But when evaluating a tool, or debugging why something didn't run, it helps to separate the actual scheduling component (deciding when) from everything built around it (what happens on success, on failure, and in what order relative to other jobs).
See how to set one up in the scheduling tutorial, keep an existing job running with the app closed at Run jobs, app closed, or check what actually happened on past runs at Job run history.
14-day free trial, no card. Schedule your first job in a minute.
No credit card. 14 days. Cancel in one click.