5 min read
Folders and Structure
Load an image or annotation folder as one structured dataset: layout detection (ImageFolder, COCO, YOLO, Pascal VOC, Hugging Face, Hive), roles, links, splits, manifests and loader code.
One folder, one dataset
Drop a folder onto the Data Hub, or use Add folder. Instead of a pile of files you get a single dataset whose Files tab shows the tree. The wizard looks at the paths and offers a layout it recognises; the Structure Manager is where you refine it.
Detected layouts
The detector reports "Detected COCO" and the like with a confidence figure. It covers:
- ImageFolder: one folder per class
- COCO: a JSON annotation file beside images
- YOLO:
images/x.jpgpaired withlabels/x.txt - Pascal VOC:
JPEGImages/x.jpgpaired withAnnotations/x.xml - Hugging Face style: files named by split
- Hive partitions:
key=valuefolders
For YOLO and VOC, confidence is the genuine fraction of files that have a partner. For COCO the path heuristic is deliberately cautious; use Verify on the Roles tab to read the annotation file and confirm. If nothing matches you get "No known layout recognized in these paths" and can map roles by hand.
The Structure Manager tabs
- Roles: tag files by glob rule, such as
images/**/*.jpgas image ortrain/**as train. Rules show unmapped files so nothing is silently ignored. - Links: pairs images with their labels or annotations (COCO, YOLO and VOC), and shows how many pairs matched. Image dimensions are read so they are available as columns.
- Manifest: a table with one row per file. Add columns by extracting values from paths with a regular expression, for example
scene_(\d+)/into a scene column. Export the manifest as CSV, JSONL or Parquet. - Splits: assign each file to train, validation or test by folder rule, random percentages with a seed (optionally grouped by a regex so related files stay together), or stratified on class labels.
You can also create empty folders inline and drag files between folders; each change is a new version in the dataset's history.
Keep samples together
A random split treats every file as a sample. For an image and its label file, use folder rules or a grouping regex so they land in the same split. Folder rules preserve image and label pairs.
Loader code
From the dataset you can preview loader code for your framework: pandas, PyTorch, tf.data or Hugging Face. The code is generated from your roles and splits; it is meant to be copied into a script, and the Studio does not run it for you.
Related
Importing data, Profiles, versions and lineage, Visual Model Builder.