8 min read
PrepFlow Transform Reference
How to build, run and save a PrepFlow graph, and a category-by-category reference of its 70-plus transforms: cleaning, encoding, scaling, imbalance, text, time series, imaging, audio and validation.
Building a graph
PrepFlow opens from the Data Hub's tools in the sidebar. The canvas has three kinds of node: Source nodes (a dataset from the Data Hub), Step nodes (one transform each) and an Output node.
- Attach one or more datasets. The Datasets panel lists what the Data Hub has loaded and can bring in more.
- Drag a transform from the palette (search it by name) onto the canvas, or add it from the outline view, and connect it to what it should read.
- Fill in the step's settings in the configuration panel. Column pickers read the real columns of the upstream node, and steps with problems list them (for example, a missing column).
- Click a node to see a live preview of the data at that point, so each step's effect is visible before you move on.
Undo and redo are available from the toolbar, and the canvas shares the selection and clipboard shortcuts described in Keyboard shortcuts. Nodes can be copied between projects.
Marking the target
Mark the prediction target column or columns. The mark feeds the steps that need a label (stratified splitting, SMOTE, target encoding and others) and becomes the Output node(s) in the Model Builder.
Running, saving and reusing
- Materialize runs the whole graph and saves the result as a new dataset in the Data Hub. The log panel records what ran. With a Jupyter or Kaggle kernel attached, the same button runs on the kernel instead ("Run on Jupyter"); see Remote compute.
- Python generates inspectable pandas and scikit-learn code and a requirements file from the graph, so you can read, run or ship what the graph does.
- Save flattens the active chain into a preprocessing artifact the Model Builder can use, so the same steps run on new data at prediction time.
- Add to pipeline puts the workflow into a pipeline as a step, so it can run on a schedule.
Steps that learn from data (scalers, encoders, imputers) must be fitted on training rows only and then replayed unchanged on validation, test and new data, or information from the test set leaks into training. Inside a pipeline, the Fit preprocessing step fits a PrepFlow workflow on the training rows only and Apply preprocessing replays it.
Cleaning
- Record Deduplication, Constant Column Removal, Row-Wise Missing Value Removal, Entirely Empty Record Pruning
- Missing Value Imputation (constant or statistic), Forward and Backward Fill, Interpolation (linear or quadratic), Datetime Forward/Backward Fill
- Outlier Detection (flag), Outlier Treatment, Outlier Clipping / Winsorizing, Percentile Classification Flag
- Equi-Width and Equi-Depth Binning
- Standardize Values, Fuzzy String Clustering, String Phonetic Fingerprinting, Normalize Values Within a Delimited Cell
- Label Noise Diagnostics
Structure and shape
Keep or remove columns, rename, reorder, cast types, format and hash columns, split and merge columns, sanitise column names, promote a row to the header, Pivot, Unpivot, Explode Delimited Column to Rows, Automated Schema Discovery, Derive Column (Formula), Aggregate / Group By, Filter rows, Sort rows.
Combining
Join, Union / Append and the As-Of Temporal Join (match each row to the nearest earlier row of another table).
Encoding
One-Hot, Ordinal / Label, Target / Mean, Weight of Evidence, Frequency / Count, Hashing, and Rare Label Grouping.
Scaling and reduction
Min-Max Scaling, Z-Score Standardization and PCA / SVD dimensionality reduction.
Class imbalance
Random Over-Sampling, Random Under-Sampling, SMOTE, SMOTE-ENN, SMOTE-Tomek and NearMiss. Apply them to training rows only.
Feature engineering
Polynomial and Interaction Features, Cluster Distance Features (K-Means), Relational Aggregation Features, Lag / Lead offsets, Rolling Window aggregations and Date/Time Component Extraction.
Text
Whitespace Trimming, Case Normalization, Regex Parsing and Extraction, Tokenization, Stopword Removal, Stemming / Lemmatization, TF-IDF Vectorization and Sentence Embedding (a Hugging Face transformer running in your browser).
Images and audio
- Images: Resize, Channel Normalization, Spatial Augmentation and a Quality / Blur check. These read images imported through the Data Hub.
- Audio: Waveform Resampling and Mel Spectrogram Extraction.
On a remote kernel, steps that read browser-only media cannot run. See the limits in Remote compute.
Splitting and data checks
- Dataset Split (Train / Val / Test), reproducible with a seed, and stratified by the target when you ask.
- Data Valuation (KNN Shapley) estimates how much each training row contributes.
- Assert Not Null, Assert Value Range, Assert Unique Values and Feature Drift Detection check data against an expectation.
- IP Address Conversion and URL Parameter Parsing for specialised string fields.
Custom code
Custom Code runs your own Python on the data inside the graph when no built-in step fits.
How PrepFlow differs from the Transform window
The Data Hub's Transform window covers plain wrangling and compiles to SQL for larger tables. PrepFlow is for machine-learning preparation and fitted transforms, and it adds the target, splits and a graph you can reuse.