Skip to content

Track Masks Using CUTIE ​

SUMMARY

CutieTracker propagates seeded object masks through successive frames, preserving their object IDs. Use it when you need an object's pixel region over time, such as following parts on a conveyor.

Installation ​

Install telekinesis-trackers using the Trackers setup guide. The default ONNX bundle downloads on first construction. PyTorch and a CUTIE source checkout are not required for inference.

The Skill ​

python
labels, alive_fraction = tracker.execute(image_rgb)

Seed the tracker before calling execute():

python
import numpy as np
from PIL import Image
from telekinesis.trackers import CutieTracker

first_rgb = np.asarray(Image.open("frame_000.png").convert("RGB"))
next_rgb = np.asarray(Image.open("frame_001.png").convert("RGB"))
# A grayscale label image: 0 = background, 1 = object one, 2 = object two, ...
seed_labels = np.asarray(Image.open("labels.png").convert("L"))

tracker = CutieTracker(device="cpu")
tracker.seed(seed_labels, first_rgb)
labels, alive_fraction = tracker.execute(next_rgb)

if labels is not None:
    Image.fromarray(labels).save("tracked_labels.png")
print(f"Objects with nonempty masks: {alive_fraction:.0%}")

Both frames and the label image must have matching dimensions. For click-assisted seed masks, see RITM.

Constructor Parameters ​

ParameterDefaultDescription
model_dirNoneLocal bundle directory; omit to download and cache the default bundle
max_internal_size480Reduce the short side to at most this size, preserving aspect ratio; -1 disables reduction. Other values must be at least 16
threads1Positive ONNX Runtime intra-operation thread count
device"auto""auto", "cpu", "cuda", or "cuda:N"
mem_everyNonePositive interval between memory insertions; None uses the bundle setting
max_mem_framesNoneMemory-bank capacity, at least 2; None uses the bundle setting
top_kNonePositive number of memory matches used by attention; None uses the bundle setting

Session Methods ​

MethodBehavior
seed(labels, image_rgb, obj_ids=None)Start a new session and return its label image. By default, track all nonzero IDs in the seed mask; obj_ids selects the IDs explicitly
execute(image_rgb)Process the next frame and return (labels, alive_fraction)
correct(image_rgb, labels)Supply a corrected mask for an existing session and update its memory. Labels must use existing object IDs
reset_session()Clear the tracking state

labels is a uint8 array at the input frame's resolution, or None from execute() when no foreground remains. alive_fraction is the fraction of seeded objects with nonempty output masks, not a confidence score.

Inspect is_seeded, object_ids, and last_labels for session state. Call seed() again to change the object selection or start at a different resolution. correct() retains the session and does not add new IDs.

Resolution and Memory ​

The default bundle supports dynamic spatial dimensions and object counts. All objects are seeded together, and frame dimensions remain fixed within each session. Older fixed-shape bundles retain their exported shape and object-count requirements.

Frames are reduced according to max_internal_size, then padded to multiples of 16. Output labels return to the original frame size. Use CutieTracker(max_internal_size=-1) to preserve input resolution before padding; this can increase memory and computation.

The memory bank keeps the seed entry and evicts older subsequent entries when full. Increase mem_every to store frames less frequently, or reduce max_mem_frames to bound memory more tightly. These settings change the temporal information available for tracking.

On CUDA, CUTIE keeps graph features and recurrent state on the GPU using ONNX Runtime I/O binding and CuPy. Inputs and returned labels are still NumPy arrays. The implementation requests full FP32 CUDA math to limit prediction drift between CPU and GPU.

Common Errors ​

  • Calling execute() or correct() before seeding raises RuntimeError.
  • Changing frame dimensions within a session raises ValueError; reseed for the new size.
  • Empty or invalid seed labels, incompatible bundle shapes, and corrections containing new object IDs are rejected.

For camera and video workflows, see Video and Live Examples.