Trackers: Visual Tracking Skills
SUMMARY
Telekinesis Trackers follows objects and points across video frames. Initialize a mask or select query points on one frame, then process subsequent frames with a tracker that retains state between calls. RITM creates the initial object masks from clicks.
Choose a Tracker
| Component | Use it for | Initialization | Runtime |
|---|---|---|---|
| CUTIE | Propagating multiple object masks | An integer label image | ONNX Runtime, CPU or CUDA |
| TAPIR | Online or offline point tracking | Pixel coordinates | ONNX Runtime, CPU or CUDA |
| TAPNext++ | Causal point tracking | Pixel coordinates and a local model | ONNX Runtime or PyTorch, CPU or CUDA |
| RITM | Creating an object's seed mask | Positive and optional negative clicks | ONNX Runtime; examples use CPU |
CUTIE produces a label image for each frame. TAPIR and TAPNext++ produce point coordinates and visibility information. RITM operates on individual images and does not track across frames.
Installation
Use Python 3.11 or newer in an isolated environment:
python -m pip install telekinesis-trackersThe distribution is telekinesis-trackers; the Python import is telekinesis.trackers. The base package includes CPU ONNX Runtime, NumPy, Pillow, and requests. PyTorch and tapnet are separate requirements only for the TAPNext++ checkpoint backend.
Check the import without downloading or loading a model:
from telekinesis import trackers
print(trackers.CutieTracker)NVIDIA GPU Setup for ONNX
After installing the base package, replace CPU ONNX Runtime with its GPU distribution and install the CUDA libraries and CuPy:
python -m pip uninstall -y onnxruntime onnxruntime-gpu
python -m pip install "onnxruntime-gpu[cuda,cudnn]>=1.21,<1.27" "cupy-cuda12x[ctk]>=14,<15"This setup requires a CUDA 12-compatible NVIDIA driver. The runtime upper bound keeps this installation on CUDA 12. CuPy is used by CUTIE's CUDA memory attention and resizing.
ONNX Runtime package replacement
The CPU and GPU distributions share the same Python module directory. Installing the package's [gpu] extra alone installs both distributions; use the replacement commands above afterward. Because telekinesis-trackers declares the CPU distribution as a dependency, pip check reports it as missing after replacement. Upgrading or reinstalling the tracker package may restore it, requiring the replacement again.
Use device="cpu", device="cuda", or device="cuda:N" to select a device. For ONNX trackers, device="auto" selects CUDA when ONNX Runtime advertises that provider, otherwise CPU. CUTIE and TAPIR check that their sessions retain the selected CUDA provider. This does not guarantee that every graph operator executes on GPU.
Model Loading
CUTIE, RITM, and TAPIR download default bundles on first construction and reuse them from ~/.cache/telekinesis/trackers/. The asset origin is https://assets.telekinesis.ai/trackers/.
| Component | Cache directory below the tracker cache |
|---|---|
| CUTIE | cutie/cutie-dynamic-spatial |
| RITM | ritm/480x864-20 |
| TAPIR online | tapir-online/causal-2 |
| TAPIR offline | tapir-offline/offline-2-2frames |
Pass model_dir to use an existing local bundle instead. Bundles contain a manifest and the model graphs; a single arbitrary .onnx file is insufficient. For use without network access, supply a complete local bundle or populate the default cache beforehand.
TAPNext++ requires a local ONNX bundle directory or checkpoint path, passed as its first constructor argument. It loads the model on load() or seed().
Inputs and Results
All image inputs are NumPy uint8 RGB arrays shaped (H, W, 3). Convert OpenCV's BGR frames before calling a tracker. Frames must keep the same dimensions within an online tracking session.
| Input | Format |
|---|---|
| CUTIE seed labels | Integer array (H, W), with 0 for background and object IDs 1 through 255 |
| Query points and RITM clicks | (N, 2) pixel coordinates in (x, y) order, inside the input image |
| TAPIR offline video | uint8 RGB array (T, H, W, 3) |
PointTracks
TAPIR and TAPNext++ return PointTracks, with coordinates expressed in the original input image's pixels:
| Field | Shape | Meaning |
|---|---|---|
xy | (T, N, 2) | float32 positions in (x, y) order |
visibility | (T, N) | float32 visibility signal |
occlusion | (T, N) | float32 occlusion signal |
visible | (T, N) | Boolean visibility decisions |
frame_indices | (T,) | Integer frame indices, starting from zero for a new session |
Online calls return one frame (T = 1). Coordinates may leave the image bounds when a point moves out of view. TAPIR combines occlusion and tracking uncertainty to calculate visibility. TAPNext++ exposes thresholded visibility: its visibility values are 0 or 1, and occlusion is their complement.
Call reset_session() before starting an unrelated sequence, or reseed with a new frame and selection. Resetting TAPNext++ clears tracking state but keeps the model loaded; unload() releases it.
Tracking Skills
tracker.execute()TrackersTrack Masks Using CUTIE
Seed multiple object masks and propagate their labels through RGB video frames with CUTIE on CPU or CUDA.
tracker.step()TrackersTrack Points Using TAPIR
Track selected pixel locations online or through an offline clip with TAPIR, returning coordinates, visibility, and occlusion.
tracker.step()TrackersTrack Points Using TAPNext++
Run causal TAPNext++ point tracking from a local ONNX bundle or PyTorch checkpoint, with explicit model and session lifecycle control.
initializer.predict()TrackersInitialize Masks Using RITM
Create and refine an object mask from positive and negative clicks, then use it to seed CUTIE tracking.
For command-line workflows, see Video and Live Examples.