Skip to content

Initialize Masks Using RITM ​

SUMMARY

MaskInitializer creates one object's mask from positive and negative pixel clicks. Use it to initialize or refine a mask before passing it to CUTIE. Calls are stateless; pass the previous probability map explicitly when refining the same object.

Installation ​

Install the base package using the Trackers setup guide. The default RITM bundle downloads on first construction and resizes images internally. These examples select CPU explicitly.

The Skill ​

python
mask = initializer.predict(image_rgb, positive_points, negative_points)

Create and Refine a Mask ​

python
import numpy as np
from PIL import Image
from telekinesis.trackers import MaskInitializer

image_rgb = np.asarray(Image.open("frame_000.png").convert("RGB"))
height, width = image_rgb.shape[:2]
# Replace these with a click inside the object and a click on unwanted background.
positive = [[0.5 * (width - 1), 0.5 * (height - 1)]]
negative = [[0.1 * (width - 1), 0.1 * (height - 1)]]

initializer = MaskInitializer(device="cpu")
probability = initializer.predict_proba(image_rgb, positive_points=positive)
mask = initializer.predict(
    image_rgb,
    positive_points=positive,
    negative_points=negative,
    previous_mask=probability,
)
Image.fromarray(mask.astype(np.uint8) * 255).save("object_mask.png")

Both click lists use (x, y) coordinates in the original image. At least one positive click is required. The default bundle supports up to 20 positive and 20 negative clicks per call. Keep each object's click lists and probability map separate.

Constructor Parameters ​

ParameterDefaultDescription
model_dirNoneLocal bundle; omit to download the default RITM bundle
threshold0.5Value in [0, 1]; predict() selects pixels whose probability is greater than this value
threads1Positive ONNX Runtime intra-operation thread count
device"auto"ONNX provider selection: "auto", "cpu", "cuda", or "cuda:N". The examples use "cpu"

Prediction Parameters and Results ​

Both methods accept image_rgb, positive_points, optional negative_points=(), and keyword-only previous_mask=None.

Method or inputFormat
predict_proba(...)Returns float32 probabilities (H, W) in [0, 1]
predict(...)Returns a boolean mask (H, W)
previous_maskProbability map from the previous refinement, matching the original image dimensions, with finite values in [0, 1]

Images and previous probabilities are resized to the bundle's processing dimensions; results are resized back to the original image size. Click coordinates must be finite and inside the image. Invalid coordinates, excess clicks, or an invalid previous probability map raise ValueError.

Seed CUTIE with the Result ​

Using image_rgb and mask from the example above:

python
from telekinesis.trackers import CutieTracker

labels = np.zeros(mask.shape, dtype=np.uint8)
labels[mask] = 1
if not mask.any():
    raise ValueError("Refine the clicks before seeding: the object mask is empty.")

tracker = CutieTracker(device="cpu")
tracker.seed(labels, image_rgb)
# For each subsequent RGB frame: labels, alive_fraction = tracker.execute(frame)

For multiple objects, initialize each mask separately and assign a unique label ID to each. Resolve overlaps before calling seed(): each pixel can belong to only one object. Start an independent object without a previous_mask; the initializer has no tracking session to reset.

See live click-assisted tracking for the interactive camera workflow.