All tutorials
    Data Hub
    Beginner
    20 min

    Explore and Profile a Dataset in the Data Hub

    Import a CSV, preview it, profile every column, assign ML roles, add constraints and check for target leakage, all without leaving the browser.

    What you will do

    Load a small table, read its profile, tell DLWAY what each column means, and set rules that every future version of the data must obey. You will use the bundled Iris sample, but any CSV with a label column works.

    Before you start

    • A project open in the Studio (any template will do)
    • A CSV file. The Iris sample has five columns: four measurements and a species label.

    Step 1: Add the data

    Open the Data Hub and choose Add data. Pick Files, then choose your CSV. On the file step, look at the parse options: the delimiter, whether the first row is a header and the text encoding. For a normal CSV the defaults are right. Confirm.

    The dataset opens on its preview. Nothing has been profiled yet, and nothing has left your machine.

    Step 2: Profile it

    Choose Profile dataset. When it finishes, scroll through the columns. For each one you get its type, the number of missing values, the exact number of distinct values and a distribution. The Correlated columns section lists the strongest relationships. Use the filter columns box to find one by name.

    Make a note of anything surprising: a numeric column stored as text, a column that is almost entirely empty, or a column with only one value.

    Step 3: Assign roles

    Open the dataset's schema editor. Each column has an ML role: target, feature, ignore, id, group, timestamp or unresolved. DLWAY has already suggested roles from the names and statistics. Use Apply every suggested ML role, then check them. Set species to target and make sure the four measurements are features.

    Roles describe your intent, so they belong to the dataset and carry over to every new version.

    Step 4: Add constraints

    In the schema editor add rules the data must satisfy:

    1. A range on sepal_length, for example a minimum of 4 and a maximum of 8.
    2. A not-null rule on species.

    Each constraint shows how many rows break it. Try tightening the range to 5 to 7 and watch the count of violations change, then put it back.

    Step 5: Check for leakage

    With a target set, the schema editor flags any feature that predicts it suspiciously well. For Iris the petal measurements are strong predictors but legitimate. A column derived from the answer after the fact, such as a numeric copy of the label, is exactly what this check exists to catch, so read the warning before you train on a suspicious column.

    Step 6: Look at versions and lineage

    Open the Versions tab. A new import is v1. Open the Lineage tab to see where the dataset came from. You will use both in the next tutorials: Clean a Dataset with the Transform Window adds a version, and constraints are enforced by pipelines.

    Where to go next

    Try it in DLWΛY

    Open the Studio and follow along in a real project. There is nothing to install.

    Open Studio