4 min read
Data Hub
Load files and folders, connect live databases, and import datasets from Kaggle and Hugging Face into one versioned library.
Connecting data
Open individual files, an entire folder, or connect a live database (Postgres, MySQL, MongoDB, Snowflake, BigQuery, and more). Dozens of formats are supported, including CSV, Parquet, Arrow, COCO/YOLO annotations, and DICOM.
Dataset hubs
Search Kaggle and Hugging Face from inside the Data Hub and pull a dataset straight into your project, without downloading and re-uploading files by hand.
Preview and filter
Datasets render in a responsive, virtualized grid. Parsing happens in a dedicated Web Worker so large files don't block the interface, and SQL runs locally on DuckDB compiled to WebAssembly.
Profiles, versions and lineage
Each dataset keeps a profile of its columns, a history of versions, and a lineage graph showing what it was derived from, so a change can be traced or undone.
Using data elsewhere
Once a dataset is loaded, it becomes available to PrepFlow, to the Model Builder for training, and to the Dashboard for chart building, with no re-uploading.
Go deeper
- Importing data: formats, parse options, duplicates and limits
- Database connections and Kaggle and Hugging Face
- Profiles, versions and lineage
- The Transform window, Folders and structure and The Data Model
- Tutorial: Explore and Profile a Dataset