Skip to content

Best Practices ​

SUMMARY

Use these recommendations as starting points for reliable Iris training. Measure validation results, change one variable at a time, and adapt the settings to your dataset, model family, and available hardware.

Dataset ​

  • Keep training and validation annotations consistent, especially category IDs and names.
  • Include varied backgrounds, viewpoints, lighting conditions, scales, and object poses that represent deployment conditions.
  • Check boxes and masks visually before starting a long run.
  • Ensure every annotated object has a valid polygon or RLE mask when training an instance-segmentation model.
  • Avoid near-duplicate images across training and validation splits.

See Prepare a Dataset for supported layouts and the dataset contract.

Model Configuration ​

  • Start with a smaller model or default resolution while validating a new dataset and pipeline.
  • Use pretrained weights unless you have a specific reason to train from scratch.
  • Keep category names stable and descriptive for SAM3-LoRA because they become text prompts.
  • Change LoRA targets deliberately; adapting more components increases memory use and trainable parameters.
  • Preserve the model configuration when resuming a checkpoint.

Trainer ​

Treat the values in the training guides as starting points rather than universal defaults.

SettingGuidance
batch_sizeUse the largest value that fits reliably in GPU memory. SAM3-LoRA commonly starts at 1.
learning_rateReduce it when loss is unstable or diverges; increase cautiously when optimization is consistently too slow.
epochsTrain until validation metrics stop improving. More epochs do not guarantee better generalization.
weight_decayUse it to regularize training; tune it together with the learning rate.
gradient_clip_normLower it when gradients are unstable. SAM3-LoRA examples start at 1.0.
mixed_precisionKeep enabled on CUDA unless numerical instability requires full precision.
num_workersIncrease it when data loading limits GPU utilization; reduce it if worker processes exhaust memory.
evaluation_intervalEvaluate frequently while validating a setup, then reduce the frequency if evaluation is expensive.
checkpoint_intervalSave often enough to limit lost work without creating unnecessary storage overhead.
seedKeep fixed when comparing configurations so results are easier to interpret.

Monitor both training loss and validation metrics. Change one major parameter at a time and record the configuration associated with each run.

Validation and Checkpoints ​

  • Select models using validation metrics rather than training loss alone.
  • Use best.pt for the best monitored result and latest.pt when continuing an interrupted run.
  • Test exported artifacts on representative images before integrating them into a deployment pipeline.
  • Keep the model family, variant, resolution, class configuration, and prompt mapping compatible when resuming.
  • Retain important checkpoints outside temporary training directories.

Reproducibility ​

  • Record the Iris and PyTorch versions, model configuration, Trainer options, and dataset revision for each run.
  • Keep the random seed fixed for controlled comparisons.
  • Use TensorBoard logs to compare loss, learning rate, validation metrics, and prediction previews.
  • Store the exact category mapping alongside the dataset and exported artifact.