All documentation pages

    4 min read

    PrepFlow

    PrepFlow is DLWAY's visual data-preparation engine: build a readable graph of cleaning, encoding, scaling and resampling steps without writing pandas.

    What PrepFlow is

    PrepFlow turns data cleaning into a graph you can read from left to right: a source, a chain of steps, and an output. It opens from the Data Hub's tools once a dataset is loaded.

    On the canvas, the Select tool (V) draws a selection box when you drag on empty space, so several steps can be moved or deleted together. The Hand (H) moves the canvas instead, as does holding Space. Ctrl C, Ctrl X and Ctrl V copy, cut and paste the selected nodes, in this workflow or into another project's; a copied source is pasted as the same dataset when that project has it.

    What the steps cover

    • Cleaning: dropping duplicates, imputing missing values, handling and winsorizing outliers
    • Encoding: one-hot, ordinal, target, weight-of-evidence and hashing encoders
    • Scaling and reduction: min-max and standard scaling, PCA
    • Class imbalance: over- and under-sampling, SMOTE and its variants
    • Splitting: train/test splits you can reproduce

    Seeing the effect of each step

    Preview the data as it flows through the graph, so a mistake is caught while you are still preparing instead of after a wasted training run.

    Keeping the result

    Materialize a flow's output as a new dataset in the Data Hub, or reuse the whole flow as a step in a pipeline.

    Go deeper

    See the transform reference for every step by category, or follow Prepare Training Data with PrepFlow.