Skip to content

Video and Live Examples ​

These commands use the examples/ scripts in a trackers source checkout. Run them from that repository's root. Installing the package alone does not provide these script paths.

Setup ​

Use Python 3.11 or newer and install the checkout with its example dependencies:

sh
python -m pip install ".[examples]"

This adds OpenCV and Loguru. Interactive selection requires a graphical desktop; live tracking also requires an accessible camera. The commands below select CPU explicitly for the tracker. For CUDA, follow the GPU setup instructions after installation and use --device cuda.

Track Masks Through a Video ​

Prepare one binary mask image per object on the video's first frame. Nonzero pixels identify that object. This example uses a 640 by 480 video and masks:

sh
python examples/track_video.py --tracker cutie --video clip.mp4 --masks object_a.png object_b.png --resize 640 480 --device cpu --out tracked.mp4 --save-masks

Each mask must be exactly the selected height and width. IDs are assigned in command-line order, starting at 1; later masks overwrite earlier ones at overlapping pixels. The script resizes video frames to --resize, so choose dimensions with the intended aspect ratio and prepare masks for that resized frame.

--out selects the overlay video. Omit it to write <input_stem>_tracked.mp4. --save-masks also writes per-frame label PNGs. Use --model-dir to supply a local CUTIE bundle instead of downloading the default.

Live Mask Tracking ​

Start CUTIE on camera zero:

sh
python examples/track_live.py --tracker cutie --camera 0 --device cpu --out live_tracked.mp4

On the initial frame, left-drag to paint an object, right-drag to erase, and press Enter to start tracking. Press q during tracking to stop.

KeyAction during initialization
1 through 9Select an object ID
n / pSelect next / previous object ID
[ / ]Shrink / grow the painting brush
u / cUndo / clear the selected object
Enter or sAccept the masks and seed tracking
EscapeCancel initialization

To initialize with RITM clicks instead of painting:

sh
python examples/track_live.py --tracker cutie --camera 0 --ritm --device cpu

Left-click inside the selected object and right-click on unwanted regions. Switch object IDs to initialize more objects, then press Enter. --initializer-dir supplies a local RITM bundle and also enables click initialization. The live script constructs RITM with its default device="auto"; --device controls CUTIE only. To select RITM's device explicitly, use the Python initializer API.

With the default dynamic CUTIE bundle, this script preserves camera resolution and defaults to --max-internal-size -1. The CutieTracker Python constructor instead defaults to 480. Set --max-internal-size 480 to reduce internal computation, or --resize WIDTH HEIGHT to resize the incoming frames. Live resizing defaults to --resize-mode contain, preserving aspect ratio with padding; cover crops, and stretch changes aspect ratio.

Interactive Point Tracking ​

Select two points for the default online TAPIR bundle:

sh
python examples/track_points.py --tracker tapir --video clip.mp4 --device cpu --out tracked_points.mp4

Left-click to add points, right-click or press u to undo, then press Enter. Use --query-frame 10 to select points on a later frame; causal tracking starts there and does not produce predictions for earlier frames.

Supply coordinates to skip the picker:

sh
python examples/track_points.py --tracker tapir --video clip.mp4 --points 100,80 240,160 --device cpu --out tracked_points.mp4

Coordinates must lie inside the source video. The script writes both an overlay .mp4 and a compressed .npz with xy, visible, visibility, occlusion, and frame_indices. The default TAPIR online bundle requires exactly two points; the default offline bundle also limits the clip to two frames. See TAPIR for the offline Python API.

TAPNext++ Point Tracking ​

Supply an existing ONNX bundle:

sh
python examples/track_points.py --tracker tapnextpp --model-dir bundles/tapnextpp/512-dynamic --video clip.mp4 --device cpu --out tapnextpp_points.mp4

Or use a checkpoint after installing its backend requirements:

sh
python examples/track_points.py --tracker tapnextpp --checkpoint tapnextpp_512.ckpt --video clip.mp4 --device cuda --out tapnextpp_points.mp4

Select one or more points with the picker for a dynamic bundle or checkpoint, or pass --points. The number remains fixed during that tracking session. The point script loads the complete input video into memory even when inference is causal; use the Python seed() / step() API for a stream processed frame by frame.