Some questions aren't a single query. Flow Books let you pull a Databricks result set and run Python on it in the same file, instead of exporting a CSV to a separate notebook.
No credit card. 14 days. Cancel in one click.
Quick answer: Flow Books are notebooks that mix SQL and Python cells against a connected Databricks SQL warehouse. Run a query in one cell, pass the result into a pandas dataframe in the next, and chart it, all in the same native file. No CSV export step, no separate notebook server to keep alive.
Databricks has its own excellent notebooks, but they run against a cluster, in the browser, tied to your workspace. If what you actually want is to pull a quick result from your SQL warehouse and do a bit of Python locally, spinning up a cluster and a browser notebook is a lot of infrastructure for a five-minute question. Flow Books skip that: SQL against your warehouse, Python on the result, in a native file on your Mac.
1. Open a Flow Book against your connected Databricks warehouse. 2. Add a SQL cell and run a query; the result becomes a dataframe. 3. Add a Python cell below it and work with that dataframe directly.
Pull thirty days of pipeline run durations, then flag the outliers in Python:
-- SQL cell SELECT run_id, pipeline_name, duration_seconds FROM ops.gold.pipeline_runs WHERE run_date >= current_date() - interval 30 days;
# Python cell
threshold = df["duration_seconds"].quantile(0.95)
outliers = df[df["duration_seconds"] > threshold]
outliers.sort_values("duration_seconds", ascending=False)
The dataframe the SQL cell produced feeds straight into the Python cell, no export, no separate kernel to keep track of.
If the answer is really just the query result, use the SQL editor directly, it's faster than opening a notebook for a single SELECT. Flow Books earn their keep on the questions that need a second step.
Flow Books are part of Studio. Scheduling a query's output or syncing results into another warehouse needs Pipelines. See pricing.
If a column gets renamed or dropped upstream, the SQL cell fails the same way a standalone query would, with a clear error rather than a silently wrong dataframe. A Flow Book isn't more fragile than a plain query against the same table, but it isn't more forgiving either, so treat schema changes the same way you would anywhere else.
Say a query pulled pipeline run durations and you want to flag which pipelines are on a known-flaky list, not something tracked in the warehouse itself:
# Python cell, after the SQL cell above flaky = set(["ingest_orders", "sync_inventory"]) df["known_flaky"] = df["pipeline_name"].isin(flaky) df[df["known_flaky"]]
That's the kind of join SQL alone handles awkwardly when one side isn't in the warehouse, and it's exactly what the Python cell is for.
Label your SQL and Python cells with something more useful than default numbering if the Flow Book will outlive the afternoon you wrote it. A short description above each cell costs a few seconds now and saves you from re-reading the whole notebook the next time you open it.
A Flow Book can live next to the plain SQL queries you've saved from the editor; there's no need to standardize on one workflow. Use the standalone editor for a query you'll run as-is, and reach for a Flow Book once a question genuinely needs a Python step on top.
The same 100,000-row and 25 MB limits that apply to a plain query apply to a SQL cell's result too. For a larger analysis, aggregate down in the SQL cell rather than trying to pull raw rows past that limit into Python.
Save the file somewhere shared and note in the first cell which connection it expects. Someone opening it later needs their own working Databricks connection with the same access; the notebook itself doesn't carry your credentials along with it.
Treat a shared Flow Book the way you'd treat any other shared analysis file, keep a copy in whatever your team already uses for version history, rather than relying on a single unversioned file that anyone can overwrite silently.
No, it's a different tool for a different moment: quick, local analysis against your SQL warehouse, not cluster-based data engineering work.
No, Flow Books query your existing SQL warehouse; there's no separate compute to spin up for the notebook.
Yes, the same catalog.schema.table structure as the standalone SQL editor.
Common data libraries are available; check the in-app docs for the current list.
14-day free trial, no card. Open a Flow Book against your Databricks warehouse.
No credit card. 14 days. Cancel in one click.