Skip to content

Track Points Using TAPIR ​

SUMMARY

TapirTracker follows selected pixel locations. Use online mode for sequential frames, or offline mode to process a complete clip using a compatible bundle.

Installation ​

Follow the Trackers setup guide. TAPIR uses ONNX Runtime on CPU or CUDA; no PyTorch installation is needed for inference.

Default bundle dimensions

Both default bundles require exactly two query points. The offline bundle also requires exactly two frames. Online mode accepts successive frames through step(). Changing mode does not make these dimensions dynamic; other counts require a compatible custom bundle.

The Skill ​

python
tracks = tracker.step(image_rgb)

Online Tracking ​

python
import numpy as np
from PIL import Image
from telekinesis.trackers import TapirTracker

first_rgb = np.asarray(Image.open("frame_000.png").convert("RGB"))
next_rgb = np.asarray(Image.open("frame_001.png").convert("RGB"))
height, width = first_rgb.shape[:2]
points = np.array([
    [0.25 * (width - 1), 0.5 * (height - 1)],
    [0.75 * (width - 1), 0.5 * (height - 1)],
], dtype=np.float32)

tracker = TapirTracker(mode="online", device="cpu")
first_tracks = tracker.seed(first_rgb, points)
tracks = tracker.step(next_rgb)
print(tracks.xy[0])       # (2, 2): one (x, y) position per point
print(tracks.visible[0])  # (2,): visibility decisions

Choose points on features you want to follow; the fractional coordinates above illustrate the input format. seed() returns predictions for frame zero, and each step() advances one frame. Subsequent frames must match the seed frame's dimensions.

Offline Tracking ​

Using the two frames and points above:

python
tracker = TapirTracker(mode="offline", device="cpu")
video_rgb = np.stack([first_rgb, next_rgb])
tracks = tracker.track(video_rgb, points, query_frame=0)
print(tracks.xy.shape)  # (2, 2, 2): frames, points, coordinates

query_frame identifies the frame on which the supplied points were selected. It must index a frame in the input clip.

Constructor Parameters ​

ParameterDefaultDescription
model_dirNoneLocal ONNX bundle; omit to download the default for the requested mode
modeNoneDefaults to online for pretrained loading, or infers the mode from a local bundle. Accepts "online", "causal", and "offline"
model_urlNoneOptional custom bundle download URL; requires model_dir
visibility_threshold0.5Threshold in [0, 1]; a point is visible when its visibility exceeds this value
threads1Positive ONNX Runtime intra-operation thread count
device"auto""auto", "cpu", "cuda", or "cuda:N"

An explicit mode must match the bundle. The mode attribute preserves the manifest spelling, "causal" or "offline"; "online" is an input alias for "causal".

Results and Session State ​

All tracking methods return PointTracks. Internal processing uses 256 by 256 images, but output coordinates are mapped back to the original image size. visibility combines the probabilities of being unoccluded and reliably tracked; occlusion reports occlusion alone.

last_tracks holds the latest result. reset_session() clears cached features, recurrent state, and frame numbering. Reseed to start another online sequence or select different points.

Common Errors ​

  • step() requires an already seeded online session.
  • track() requires an offline bundle; seed() and step() require a causal bundle.
  • Query-point counts and offline frame counts must match the bundle manifest.
  • Query coordinates must be finite and inside the query image; online frames must retain their original size.

See interactive point selection for a video example.