Best Practices
SUMMARY
Use these recommendations as starting points for reliable Iris training. Measure validation results, change one variable at a time, and adapt the settings to your dataset, model family, and available hardware.
Dataset
- Keep training and validation annotations consistent, especially category IDs and names.
- Include varied backgrounds, viewpoints, lighting conditions, scales, and object poses that represent deployment conditions.
- Check boxes and masks visually before starting a long run.
- Ensure every annotated object has a valid polygon or RLE mask when training an instance-segmentation model.
- Avoid near-duplicate images across training and validation splits.
See Prepare a Dataset for supported layouts and the dataset contract.
Model Configuration
- Start with a smaller model or default resolution while validating a new dataset and pipeline.
- Use pretrained weights unless you have a specific reason to train from scratch.
- Keep category names stable and descriptive for SAM3-LoRA because they become text prompts.
- Change LoRA targets deliberately; adapting more components increases memory use and trainable parameters.
- Preserve the model configuration when resuming a checkpoint.
Trainer
Treat the values in the training guides as starting points rather than universal defaults.
| Setting | Guidance |
|---|---|
batch_size | Use the largest value that fits reliably in GPU memory. SAM3-LoRA commonly starts at 1. |
learning_rate | Reduce it when loss is unstable or diverges; increase cautiously when optimization is consistently too slow. |
epochs | Train until validation metrics stop improving. More epochs do not guarantee better generalization. |
weight_decay | Use it to regularize training; tune it together with the learning rate. |
gradient_clip_norm | Lower it when gradients are unstable. SAM3-LoRA examples start at 1.0. |
mixed_precision | Keep enabled on CUDA unless numerical instability requires full precision. |
num_workers | Increase it when data loading limits GPU utilization; reduce it if worker processes exhaust memory. |
evaluation_interval | Evaluate frequently while validating a setup, then reduce the frequency if evaluation is expensive. |
checkpoint_interval | Save often enough to limit lost work without creating unnecessary storage overhead. |
seed | Keep fixed when comparing configurations so results are easier to interpret. |
Monitor both training loss and validation metrics. Change one major parameter at a time and record the configuration associated with each run.
Validation and Checkpoints
- Select models using validation metrics rather than training loss alone.
- Use
best.ptfor the best monitored result andlatest.ptwhen continuing an interrupted run. - Test exported artifacts on representative images before integrating them into a deployment pipeline.
- Keep the model family, variant, resolution, class configuration, and prompt mapping compatible when resuming.
- Retain important checkpoints outside temporary training directories.
Reproducibility
- Record the Iris and PyTorch versions, model configuration, Trainer options, and dataset revision for each run.
- Keep the random seed fixed for controlled comparisons.
- Use TensorBoard logs to compare loss, learning rate, validation metrics, and prediction previews.
- Store the exact category mapping alongside the dataset and exported artifact.