Skip to content

Prepare a Dataset ​

SUMMARY

Iris trains from COCO annotations. Detection needs boxes and category IDs; RF-DETR segmentation and SAM3-LoRA also need a segmentation polygon or RLE mask for every annotated object.

Dataset Workflow ​

Load Dataset ​

Use the transforms and mask requirement supplied by the model wrapper:

python
from telekinesis.iris.dataset import COCODataset
from telekinesis.iris.models import RFDETR

model = RFDETR(variant="seg-nano", num_classes=3)

training_dataset = COCODataset(
    "dataset/train",
    transforms=model.train_transforms,
    include_masks=model.requires_masks,
)
validation_dataset = COCODataset(
    "dataset/valid",
    transforms=model.val_transforms,
    include_masks=model.requires_masks,
)

SAM3LoRA.requires_masks is always True; its adapter performs resizing and normalization, so its transform properties are None.

Inspect Categories ​

COCO category IDs become training labels. Read them before choosing num_classes or constructing SAM3 prompts:

python
metadata = COCODataset("dataset/train")
category_ids = metadata.coco.getCatIds()
categories = {
    category["id"]: category["name"]
    for category in metadata.coco.loadCats(category_ids)
}

For RF-DETR, make the classification head large enough for the greatest category ID:

python
model = RFDETR(variant="seg-nano", num_classes=max(category_ids))

For SAM3-LoRA, pass the mapping directly as its training prompts:

python
from telekinesis.iris.models import SAM3LoRA

model = SAM3LoRA(class_names=categories)

Segmentation annotations are required

When include_masks=True, every object must have a non-empty COCO segmentation. Iris raises an error rather than silently training without a mask.

Supported Layouts ​

Use the standard Iris layout:

text
dataset/train/
├── images/
│   ├── image-001.jpg
│   └── image-002.jpg
└── annotations.json

Iris also recognizes a Roboflow-style COCO split, where images and _annotations.coco.json are in the same directory:

text
dataset/train/
├── image-001.jpg
├── image-002.jpg
└── _annotations.coco.json

The annotation file must contain the usual COCO images, annotations, and categories arrays. Bounding boxes use COCO's [x, y, width, height] format in JSON; COCODataset converts them to xyxy tensors.

Dataset Reference ​

Each dataset item is (image, target). The image is an RGB float32 tensor in CHW layout and the [0, 1] range. The target contains:

KeyMeaning
boxesAbsolute xyxy boxes, shape (N, 4)
labelsCOCO category IDs
image_idCOCO image ID
areaObject areas
iscrowdCOCO crowd flags
masksOptional binary masks, shape (N, H, W)

Next Step

Continue with Train and Resume a Model.