Prepare Training Data with PrepFlow
Build a visual preparation graph for a messy table: drop useless columns, impute, encode and scale, mark the target, split, preview every step, then materialize the result and generate the equivalent Python.
What you will build
A PrepFlow graph that turns the Titanic passenger table into model-ready numeric data. The same pattern works for any tabular classification task.
Before you start
- The Titanic CSV loaded in the Data Hub (see Explore and Profile a Dataset for how to import a CSV)
- About half an hour
Step 1: Open PrepFlow with your data
In the sidebar open the Data Hub, then PrepFlow. Attach the dataset from the Datasets panel (you can load more from the Data Hub there). A Source node appears on the canvas.
Step 2: Drop columns you will not use
Search the palette for Keep / Remove Columns and drag it onto the canvas. Connect the Source to it and remove Name, Ticket and Cabin. Click the node: the live preview shows the table after this step. These columns are free text or mostly empty, so a model cannot use them directly.
Step 3: Fill missing values
Add Missing Value Imputation (Statistic) and fill Age with its median. Add Missing Value Imputation (Constant) for Embarked if it has gaps. Watch the preview: the missing counts should drop to zero for those columns.
Step 4: Encode categories
Add One-Hot Encoding for Sex and Embarked. Each category becomes its own 0/1 column. Alternatives in the palette include Ordinal / Label Encoding and Target / Mean Encoding; pick one-hot here because the categories are few and unordered.
Step 5: Scale numbers
Add Z-Score Standardization to Age and Fare. Scaling puts numeric columns on a comparable range, which helps neural networks train.
Step 6: Mark the target
Mark Survived as the prediction target. The mark feeds the steps that need a label and becomes the Output node in the Model Builder.
Step 7: Split
Add Dataset Split (Train / Val / Test), set the proportions and a seed, and stratify by the target so each partition keeps the same survival rate. The seed makes the split repeatable.
Step 8: Check each step
Click through the nodes from the Source to the end. Any step with a problem lists it, for example a column that no longer exists. Use undo and redo from the toolbar, or open the outline view to see the graph as a list.
Step 9: Materialize
Press Materialize. The log shows what ran, and the result is saved as a new dataset in the Data Hub, ready to train on. If a Jupyter or Kaggle kernel is attached, the button offers to run on it instead.
Step 10: Read and keep the code
Press Python to generate inspectable pandas and scikit-learn code with a requirements file. Read it to confirm it does what you drew. Press Save to keep the chain as a preprocessing artifact the Model Builder can use at prediction time.
A note on leakage
Scalers and imputers learn from data. If you fit them on the whole table before splitting, information from the test rows leaks into training. In a pipeline, use Fit preprocessing, which fits on training rows only and replays on the rest. See Build Your First Pipeline.
Related
PrepFlow and the transform reference.