Organise an Image and Annotation Folder
Load a folder of images and YOLO or COCO labels as one dataset, confirm the layout, link images to labels, create splits that keep pairs together and export a manifest and loader code.
What you will do
Turn a folder of images with labels into one dataset with roles, links and train/validation/test splits, ready for training code. A small public example is the COCO8 sample from Ultralytics; any folder with a similar layout works.
Before you start
- A folder with images and either YOLO label files (
labels/x.txtforimages/x.jpg) or a COCO annotation JSON - A Chromium browser to pick a folder easily
Step 1: Load the folder
In the Data Hub choose Add folder (or drop the folder). The wizard reads the paths, and the stepper shows the layout it detected, for example "YOLO detected" with a confidence. A folder with a recognised layout takes you to the Structure step.
Step 2: Confirm the layout
Open the Structure Manager. The header shows the detected layout. For COCO the path-only confidence is deliberately modest, so press Verify on the Roles tab to read the annotation file and confirm. If nothing is detected, you will see that no known layout was recognised, and you can set up roles by hand.
Step 3: Roles
On the Roles tab, check the glob rules that tag files: images, labels and annotation files. Unmapped files are marked, so nothing is silently ignored. Edit a rule, for example images/**/*.jpg, and see the file counts change.
Step 4: Links
On Links, DLWAY pairs each image with its label (or its annotations in COCO) and reports how many pairs matched. Image dimensions are read so they can be used as columns. Fix any unmatched files by correcting a rule or a name.
Step 5: Manifest
The Manifest tab is a table with one row per file. Add a column from the path with a regular expression, such as scene_(\d+)/ into a scene column. Export the manifest as CSV, JSONL or Parquet.
Step 6: Splits
On Splits choose a strategy:
- Folder rule: a path pattern such as
train/**assigns files to a split. This keeps image and label pairs together. - Random: percentages with a seed. Group by a regex so related files, like all frames of one scene, stay in one split.
- Stratified: keeps class proportions, using labels read from the annotations.
Remember that a random split treats each file as a sample; with separate image and label files, prefer a folder rule or a grouping regex.
Step 7: Version history
Create a new empty folder inline and drag a file into it. Open Versions: each structural change is a new version of the dataset.
Step 8: Loader code
Preview the loader code for pandas, PyTorch, tf.data or Hugging Face, generated from your roles and splits. Copy it into a script. The Studio shows the code but does not run it.