Prepare a Dataset
SUMMARY
Iris trains from COCO annotations. Detection needs boxes and category IDs; RF-DETR segmentation and SAM3-LoRA also need a segmentation polygon or RLE mask for every annotated object.
Dataset Workflow
Load Dataset
Use the transforms and mask requirement supplied by the model wrapper:
from telekinesis.iris.dataset import COCODataset
from telekinesis.iris.models import RFDETR
model = RFDETR(variant="seg-nano", num_classes=3)
training_dataset = COCODataset(
"dataset/train",
transforms=model.train_transforms,
include_masks=model.requires_masks,
)
validation_dataset = COCODataset(
"dataset/valid",
transforms=model.val_transforms,
include_masks=model.requires_masks,
)SAM3LoRA.requires_masks is always True; its adapter performs resizing and normalization, so its transform properties are None.
Inspect Categories
COCO category IDs become training labels. Read them before choosing num_classes or constructing SAM3 prompts:
metadata = COCODataset("dataset/train")
category_ids = metadata.coco.getCatIds()
categories = {
category["id"]: category["name"]
for category in metadata.coco.loadCats(category_ids)
}For RF-DETR, make the classification head large enough for the greatest category ID:
model = RFDETR(variant="seg-nano", num_classes=max(category_ids))For SAM3-LoRA, pass the mapping directly as its training prompts:
from telekinesis.iris.models import SAM3LoRA
model = SAM3LoRA(class_names=categories)Segmentation annotations are required
When include_masks=True, every object must have a non-empty COCO segmentation. Iris raises an error rather than silently training without a mask.
Supported Layouts
Use the standard Iris layout:
dataset/train/
├── images/
│ ├── image-001.jpg
│ └── image-002.jpg
└── annotations.jsonIris also recognizes a Roboflow-style COCO split, where images and _annotations.coco.json are in the same directory:
dataset/train/
├── image-001.jpg
├── image-002.jpg
└── _annotations.coco.jsonThe annotation file must contain the usual COCO images, annotations, and categories arrays. Bounding boxes use COCO's [x, y, width, height] format in JSON; COCODataset converts them to xyxy tensors.
Dataset Reference
Each dataset item is (image, target). The image is an RGB float32 tensor in CHW layout and the [0, 1] range. The target contains:
| Key | Meaning |
|---|---|
boxes | Absolute xyxy boxes, shape (N, 4) |
labels | COCO category IDs |
image_id | COCO image ID |
area | Object areas |
iscrowd | COCO crowd flags |
masks | Optional binary masks, shape (N, H, W) |
Next Step
Continue with Train and Resume a Model.