All tutorials
    Data Hub
    Intermediate
    30 min

    Organise an Image and Annotation Folder

    Load a folder of images and YOLO or COCO labels as one dataset, confirm the layout, link images to labels, create splits that keep pairs together and export a manifest and loader code.

    What you will do

    Turn a folder of images with labels into one dataset with roles, links and train/validation/test splits, ready for training code. A small public example is the COCO8 sample from Ultralytics; any folder with a similar layout works.

    Before you start

    • A folder with images and either YOLO label files (labels/x.txt for images/x.jpg) or a COCO annotation JSON
    • A Chromium browser to pick a folder easily

    Step 1: Load the folder

    In the Data Hub choose Add folder (or drop the folder). The wizard reads the paths, and the stepper shows the layout it detected, for example "YOLO detected" with a confidence. A folder with a recognised layout takes you to the Structure step.

    Step 2: Confirm the layout

    Open the Structure Manager. The header shows the detected layout. For COCO the path-only confidence is deliberately modest, so press Verify on the Roles tab to read the annotation file and confirm. If nothing is detected, you will see that no known layout was recognised, and you can set up roles by hand.

    Step 3: Roles

    On the Roles tab, check the glob rules that tag files: images, labels and annotation files. Unmapped files are marked, so nothing is silently ignored. Edit a rule, for example images/**/*.jpg, and see the file counts change.

    On Links, DLWAY pairs each image with its label (or its annotations in COCO) and reports how many pairs matched. Image dimensions are read so they can be used as columns. Fix any unmatched files by correcting a rule or a name.

    Step 5: Manifest

    The Manifest tab is a table with one row per file. Add a column from the path with a regular expression, such as scene_(\d+)/ into a scene column. Export the manifest as CSV, JSONL or Parquet.

    Step 6: Splits

    On Splits choose a strategy:

    • Folder rule: a path pattern such as train/** assigns files to a split. This keeps image and label pairs together.
    • Random: percentages with a seed. Group by a regex so related files, like all frames of one scene, stay in one split.
    • Stratified: keeps class proportions, using labels read from the annotations.

    Remember that a random split treats each file as a sample; with separate image and label files, prefer a folder rule or a grouping regex.

    Step 7: Version history

    Create a new empty folder inline and drag a file into it. Open Versions: each structural change is a new version of the dataset.

    Step 8: Loader code

    Preview the loader code for pandas, PyTorch, tf.data or Hugging Face, generated from your roles and splits. Copy it into a script. The Studio shows the code but does not run it.

    Folders and structure and Importing data.

    Try it in DLWΛY

    Open the Studio and follow along in a real project. There is nothing to install.

    Open Studio