All documentation pages

    4 min read

    Pipelines

    Chain reading data, preparation, training, code and reporting into a pipeline you can rerun, schedule, or trigger when a dataset changes.

    Repeatable work

    A pipeline is a set of steps that run in order: read a dataset, apply a PrepFlow graph, train or evaluate a model, run code, write a report. Build it on the canvas or in the outline view, then run all of it or just a part. On the canvas, drag on empty space to select several steps at once (V), or switch to the hand (H) to move around. Ctrl C, Ctrl X and Ctrl V copy, cut and paste the selected steps as pipeline-file text, so they can go into another pipeline or another project.

    Schedules and triggers

    A pipeline can run on a schedule or whenever one of its input datasets changes. Runs happen in your browser, so a scheduled pipeline runs while the Studio is open in a tab.

    Run history

    Each run is recorded with its inputs, outputs and status, and two runs can be compared side by side. How many runs are kept, and how often a step that could not start is retried, are set under Settings → Data and pipelines.

    Writing to a database

    A pipeline can write its results to a connected database. Before it does, it shows what it is about to change and asks for approval.

    Sharing a pipeline

    Export a pipeline as a definition file and import it into another project.

    Go deeper