7 min read
Profiles, Versions and Lineage
What the Data Hub knows about a dataset: its column profile and quality checks, schema roles and constraints, version history with diff and rollback, a lineage graph, storage usage and portable bundles.
The library
The Data Hub opens on a dataset library. Search by dataset or column name (Ctrl+K), switch between grid and list, sort, tag datasets, group them into collections, and archive or delete in bulk. Archive hides a dataset and keeps its history; deleted datasets go to a trash from which they can be restored. Smart collections filter on properties such as format or tags.
Opening a dataset shows tabs: Preview, Columns, Files (for folder datasets), Versions and Lineage, with the schema editor one click away.
Profile
Profiling is optional and starts when you choose Profile dataset. It reports for every column the type, the number of missing values, the exact number of distinct values and the distribution, plus data-quality findings and the strongest correlated columns. Results are saved per dataset version and reused until the data changes.
How much is shown automatically when you open a dataset is a setting: Settings → Data and pipelines → Data Analysis can be Off, Quick (compact charts and key statistics) or Complete (full statistics and larger charts).
If a snapshot holds only the first rows of a bigger source, the profile says so ("a partial copy").
Schema, roles and constraints
The schema editor records what each column means, separately from its storage type.
- ML roles: target, feature, ignore, id, group, timestamp or unresolved. DLWAY suggests roles from column names and statistics (for example, an
_idcolumn becomes id, a column calledlabelbecomes the target). Apply every suggestion with one click or override each one. Roles belong to the dataset, not a version, because they describe intent. - Type changes: preview a conversion before applying it. The preview reports values that would be out of range.
- Constraints: a range (min and/or max), a uniqueness rule or a not-null rule per column. Each is re-checked against every new version, and the editor shows the violation count.
- Leakage check: once a target is set, features that predict it suspiciously well are flagged. For a categorical target it measures how often the majority class within each feature value is right; for a numeric target it uses correlation.
- Hashing a column replaces values with their one-way SHA-256 hash, useful for removing identifiers.
Constraints saved here are also what the Validate data pipeline step checks.
Versions
Every dataset has a version history. A new import starts at v1, and each transform, re-import or pipeline output adds a version. The Versions tab lists them with who made them and how, lets you compare any two with a diff, and offers three actions:
- Promote a newer version to be the current one.
- Roll back by making an older version current again; nothing is deleted.
- Pin for training to record the version training should read. A new version never replaces a pin silently.
Lineage
The Lineage tab draws what a dataset came from and what was built from it. Dataset level shows datasets, recipes, models, dashboards and pipelines; Column level traces individual columns through recipes. Stale dependents, items built from a version that has since been replaced, are highlighted so you know what to refresh.
Storage
Local Storage shows how much of your browser's quota the project uses and lists every stored file. You can delete a file, remove files that no dataset references any more, view the background job history, and reconcile the newer per-project transform storage. Cleanup shows what it will remove and keeps the content that current and historical versions still need.
Bundles: moving datasets between projects
Export bundle packages datasets and their versions into one archive, including parsing choices, so an imported copy keeps its version numbers and profile settings. Choose which versions to include (a bundle holds at most 4 GB) and where to save: Downloads (built in memory, so very large bundles may fail) or a project folder (streams to disk). Importing a bundle restores the datasets in another project.
For folder datasets you can also export a manifest (one row per file) as CSV, JSONL or Parquet. See Folders and structure.