4 min read
Pipelines
Chain reading data, preparation, training, code and reporting into a pipeline you can rerun, schedule, or trigger when a dataset changes.
Repeatable work
A pipeline is a set of steps that run in order: read a dataset, apply a PrepFlow graph, train or evaluate a model, run code, write a report. Build it on the canvas or in the outline view, then run all of it or just a part. On the canvas, drag on empty space to select several steps at once (V), or switch to the hand (H) to move around. Ctrl C, Ctrl X and Ctrl V copy, cut and paste the selected steps as pipeline-file text, so they can go into another pipeline or another project.
Schedules and triggers
A pipeline can run on a schedule or whenever one of its input datasets changes. Runs happen in your browser, so a scheduled pipeline runs while the Studio is open in a tab.
Run history
Each run is recorded with its inputs, outputs and status, and two runs can be compared side by side. How many runs are kept, and how often a step that could not start is retried, are set under Settings → Data and pipelines.
Writing to a database
A pipeline can write its results to a connected database. Before it does, it shows what it is about to change and asks for approval.
Sharing a pipeline
Export a pipeline as a definition file and import it into another project.
Go deeper
- Pipeline step reference
- Running pipelines: publishing, parameters, schedules, backfill and sweeps
- Tutorial: Build Your First Pipeline