8 min read
Running Pipelines
Drafts and published revisions, parameters, partial runs, schedules and dataset triggers, backfills, parameter sweeps, run history and comparison, retries, and the settings that govern them.
The pipeline library
Pipelines opens on a library of the project's pipelines, searchable and filterable by tag. From it you can create a pipeline, start from a template, import a pipeline file, put pipelines away in the Archived view (archived pipelines stop running on a schedule and nothing is deleted), and manage the tables that pipelines keep up to date.
Templates build a ready-made draft from the assets you pick: clean data then refresh a dashboard, retrain and evaluate a model, score new data with a saved model, train-gate-and-package, keep a table up to date from a source, load only what changed from a database, check a dataset before anything uses it, summarise with SQL then refresh a dashboard, train a challenger and compare it with the current model, and export a cleaned dataset to the project folder. A template never invents an asset and never embeds a credential.
The editor
The editor shows the graph as a canvas or as an outline, with undo and redo. A palette offers Steps, and your project's Assets can be dragged straight on to the canvas. A node's configuration opens beside it, and each step summarises its settings on the card. Problems are listed rather than hidden: a missing asset (a dataset or model that has been deleted) offers replacement choices, and a banner tells you when a step is pinned to something that has since changed.
The pipeline is also available as plain text in the Definition editor, so it can be diffed, reviewed and shared.
Draft, publish, run
- Draft: your working copy, saved as you go with the canvas positions.
- Publish: turns the draft into a numbered, immutable revision. Publishing when nothing executable changed does not create a new one.
- Publish & Run: publishes and starts a run. Runs always use a published revision, so what ran is always recoverable.
Revision history lists revisions and shows a graph diff between any two. A nested Run Pipeline step uses another pipeline's published version.
Parameters
Parameters are named values you can change per run without editing the pipeline: a threshold, a dataset to read, a date window. Give each a name (letters, digits and underscores, starting with a letter), a type and a default. Bind a parameter to a step setting; a binding can only replace a setting that already exists with a value of the same kind, so a typo fails loudly instead of being ignored. A parameter is a value, never spliced into SQL, a shell command or code, and the settings that decide what code runs cannot be bound.
Running
Before a run, an estimate panel states what can be measured: which datasets are read (version, rows, size), which steps use TensorFlow.js or run your own code, and a warning when a source has more rows than a later step accepts. It does not guess times.
- Start from step runs just part of a pipeline, reusing earlier results.
- The run timeline shows each step as it starts, succeeds, fails, is skipped (for example "branch not taken") or is withheld by a quality gate.
- Run details list each step's inputs, outputs and messages, including the datasets and models it produced, which are opened in their own modules.
- A step that failed in a way that can pass on its own (training, evaluation and similar) is retried with a growing wait. A step that runs your code is never retried.
- A run that writes to a database stops and shows what it is about to change, according to that connection's write mode. See Database connections.
- If you try to switch or close the project while a pipeline is running, the Studio asks whether to stop the run first.
Runs execute in your browser tab. They continue while the Studio is open and stop if the tab closes. A late result from a tab that lost its claim on the project cannot commit.
Schedules and triggers
Open a pipeline's Schedule section to run it automatically:
- Timed: every minute, 5, 15 or 30 minutes, hour, 6 hours or day; daily at a time you choose; or a cron expression (the editor previews the next time it will fire).
- When a dataset gets a new version: pick the dataset to watch. Nothing runs for data that was already there when you turned it on.
- Give automatic runs their own parameter values; blank means the default.
- Missed runs: if the project was closed when a run came due, a new schedule either runs once to catch up or skips. The default for new schedules is set in Settings.
Because runs happen in the browser, a scheduled pipeline runs only while the Studio is open in a tab.
Backfill
Backfill runs a pipeline once per window of a past period, oldest first: a number of days, weeks or months between a first date and a date to stop before (up to 60 windows). Each window is a normal run whose "from" and "to" parameters are its two ends, so a step's filter decides what a window means. Windows are half-open, so consecutive windows meet without overlapping, and dates are whole calendar dates counted in UTC. Pair it with Replace Partitions so re-running a window gives the same table.
Parameter sweeps
A sweep runs the same published revision once per combination of values you list for one or more parameters (up to 12 runs), one after another. Each is a normal recorded run, steps that do not depend on the swept values are reused, and the results can be compared like any other runs. A pipeline with a New Rows step finds nothing new on the second run of a sweep, because the first run moved its position forward.
History and comparison
Every run is recorded with its inputs, outputs, status and parameter values. Compare two or more runs side by side: parameters, evaluation metrics, a model's final training numbers and a quality gate's verdict, with the better value of a metric such as accuracy or loss marked. How many runs are kept is a setting, and Clean up history shows what it would remove before you confirm.
Sharing and keeping files
- Export a pipeline as a definition file and import it in another project. Steps copied with Ctrl+C are plain pipeline-file text and paste into another pipeline.
- Project folder copies your pipelines into the connected project folder and can restore them from it.
- Tables kept up to date by pipelines lists the datasets that Upsert and Replace Partitions maintain. A table with many versions can be compacted, keeping only the newest ones you choose.
Settings
Under Settings → Data and pipelines:
- Runs kept per pipeline: default 50 (5 to 200). Older finished runs are removed when a new one starts; runs in progress, and the run an incremental step needs, are always kept.
- "Clean up history" keeps: default 10 runs (1 to 100).
- Tries for a step that could not start: default 3 (1 to 5).
- First wait before a retry: default 2 seconds (1 to 60); each wait doubles, up to half a minute.
- Missed runs for new schedules: run once, or skip.
Related
Pipeline step reference, Jobs and notifications, Database connections.