Skip to content

Estimate Pose Using FoundationPose ​

SUMMARY

Estimate Pose Using FoundationPose estimates a 6-DOF object pose from an RGB-D frame plus a CAD mesh of the object.

It runs NVIDIA's FoundationPose, a zero-shot pose estimation and tracking model, in one of two modes: register (pass a bbox) samples and scores many pose hypotheses to estimate a pose from scratch; track (pass previous_pose) refines a prior estimate for a new frame, much more cheaply. To track an object across frames, register once, then feed each call's tracking_pose back in as the next call's previous_pose.

Use this Skill when you have a CAD mesh of an object and want its 6-DOF pose for grasping, manipulation, or scene understanding, without training a per-object detector.

The Skill ​

python
from telekinesis import vitreous

# Register on the first frame.
pose, tracking_pose = vitreous.estimate_pose_using_foundation_pose(
    rgb_image=rgb_image,
    depth_image=depth_image,
    mesh=mesh,
    camera_calibration=camera_calibration,
    bbox=bbox,
)

# Track on later frames.
pose, tracking_pose = vitreous.estimate_pose_using_foundation_pose(
    rgb_image=next_rgb_image,
    depth_image=next_depth_image,
    mesh=mesh,
    camera_calibration=camera_calibration,
    previous_pose=tracking_pose,
)
API Reference
Full parameter and return type documentation for estimate_pose_using_foundation_pose.
View Reference →

The Code ​

python
"""
Demonstrates estimating a 6-DOF object pose from an RGB-D frame and a CAD
mesh using FoundationPose, then tracking it across a second frame.
"""

import numpy as np
from loguru import logger
import rerun as rr

from telekinesis import vitreous, datatypes


def estimate_pose_using_foundation_pose_example():
    """
    Registers a pose on a synthetic first frame, then tracks it on a
    second frame using the returned tracking_pose.
    """
    # ===================== Load Data ==========================================
    height, width = 480, 640
    rgb_image = np.full((height, width, 3), 200, dtype=np.uint8)
    depth_image = np.full((height, width), 1.0, dtype=np.float32)  # a flat wall 1m away

    # A small cube mesh, centered at the origin.
    mesh = datatypes.Mesh3D(
        vertex_positions=[
            [-0.05, -0.05, -0.05], [0.05, -0.05, -0.05], [0.05, 0.05, -0.05], [-0.05, 0.05, -0.05],
            [-0.05, -0.05, 0.05], [0.05, -0.05, 0.05], [0.05, 0.05, 0.05], [-0.05, 0.05, 0.05],
        ],
        triangle_indices=[
            [0, 1, 2], [0, 2, 3], [4, 6, 5], [4, 7, 6],
            [0, 4, 5], [0, 5, 1], [1, 5, 6], [1, 6, 2],
            [2, 6, 7], [2, 7, 3], [3, 7, 4], [3, 4, 0],
        ],
    )

    camera_calibration = datatypes.CameraCalibration(
        width=width,
        height=height,
        distortion_model="plumb_bob",
        distortion_parameters=[0.0, 0.0, 0.0, 0.0, 0.0],
        intrinsic_matrix=[600.0, 0.0, 320.0, 0.0, 600.0, 240.0, 0.0, 0.0, 1.0],
    )
    bbox = [280, 200, 360, 280]  # a small region roughly at the image center

    # ===================== Run Skill ==========================================
    # Frame 1: register from scratch using the bbox.
    pose, tracking_pose = vitreous.estimate_pose_using_foundation_pose(
        rgb_image=rgb_image,
        depth_image=depth_image,
        mesh=mesh,
        camera_calibration=camera_calibration,
        bbox=bbox,
    )

    # Frame 2: track forward using the previous frame's tracking_pose.
    pose, tracking_pose = vitreous.estimate_pose_using_foundation_pose(
        rgb_image=rgb_image,
        depth_image=depth_image,
        mesh=mesh,
        camera_calibration=camera_calibration,
        previous_pose=tracking_pose,
    )

    # ===================== Log ================================================
    logger.success(f"Estimated pose for {mesh}")
    logger.success(f"Results: {pose}")
    logger.info(f"Pose matrix:\n{pose.data}")

    # ===================== Visualization  (Optional) ===========================
    rr.init("estimate_pose_using_foundation_pose_example", spawn=True)
    datatypes.visualize(mesh, entity_path="/1-mesh")


if __name__ == "__main__":
    estimate_pose_using_foundation_pose_example()

Runnable examples are available in the Telekinesis examples repository.

Follow the README in that repository to set up the environment, run this specific example with:

bash
cd telekinesis-examples
python examples/point_cloud/estimate_pose_using_foundation_pose.py

Parameter Configuration ​

ParameterTypeDefaultDescription
rgb_imagedatatypes.Image | np.ndarrayrequiredThe RGB image, shape (H, W, 3).
depth_imagedatatypes.DepthImage | np.ndarrayrequiredThe metric depth map, aligned to rgb_image, shape (H, W), in meters.
meshdatatypes.Mesh3D | dictrequiredThe tracked object's CAD mesh. Only vertex_positions/triangle_indices are used.
camera_calibrationdatatypes.CameraCalibration | dictrequiredThe camera's intrinsic calibration. Only intrinsic_matrix is used.
bboxdatatypes.Box2D | np.ndarray | listNone2D bounding box [umin, vmin, umax, vmax] (XYXY) around the object, used to register a new pose. Exactly one of bbox/previous_pose must be given.
previous_posedatatypes.Transform3D | np.ndarrayNoneA prior call's tracking_pose, used to track an existing pose. Exactly one of bbox/previous_pose must be given.
refine_iterationsdatatypes.Int | int2Number of refine iterations to run.

Returns ​

TypeDescription
(pose, tracking_pose) tuple of two datatypes.Transform3Dpose is the object-to-camera transform in the mesh's own (uncentered) coordinate system - the pose your application should use. tracking_pose is opaque continuation state; pass it back as previous_pose on a later call to track this object more cheaply than re-registering. Don't interpret its values directly.

Raises ​

ExceptionCondition
TypeErrorA parameter's value does not match its expected type (see the Parameter Configuration table above)
ValueErrorNeither or both of bbox/previous_pose are given
ConfigurationErrorThe TELEKINESIS_API_KEY environment variable is not set
SerializationErrorThe request input failed to serialize, the response was not returned as an Arrow stream, or the response failed to deserialize
RequestTimeoutErrorThe request to the Vitreous service timed out
TransportErrorA network failure occurred before a response was received
ClientErrorThe Vitreous service rejected the request due to invalid or malformed input (HTTP 400/422), an unrecognized endpoint (HTTP 404), or another unexpected 4xx response
AuthenticationErrorThe API key was rejected as invalid or expired (HTTP 401)
AuthenticationServiceErrorThe authentication service returned an invalid response, was temporarily unavailable, or timed out (HTTP 502/503/504)
ServerErrorThe Vitreous service returned a 5xx or otherwise unexpected error response

How to Tune the Parameters ​

  • bbox vs. previous_pose is the main mode switch: use bbox on the first frame (or whenever you've lost track of the object), and previous_pose on subsequent frames of the same object to track it much more cheaply than re-registering.
  • refine_iterations controls how many refinement passes run per call, for both register and track. More iterations can marginally improve pose accuracy at the cost of compute time.
  • bbox should tightly enclose the object when registering - a loose or offset box widens the search and can lower accuracy or cause registration to fail.
  • Track only holds up if the object hasn't moved far between frames relative to your capture rate; if tracking drifts or loses the object, re-register with a fresh bbox.

Where to Use the Skill ​

Common pipelines include:

  • Robotic grasping – getting an object's 6-DOF pose to plan a grasp
  • Bin picking – registering pose for an object identified by a detector, then tracking it through a pick-and-place motion
  • Scene understanding – localizing known CAD objects in a scene for downstream planning

Alternative Skills ​

Skillvs. Estimate Pose Using FoundationPose
estimate_depth_using_fast_foundation_stereoProduces a metric depth map from a stereo pair, rather than an object pose. Useful as an input source for depth_image if you don't already have a depth sensor.

When Not to Use the Skill ​

Do not use Estimate Pose Using FoundationPose when:

  • You don't have a CAD mesh of the object – this Skill requires a mesh; if you only need to detect/segment the object (not its 3D pose), use a Cornea or Retina Skill instead
  • You need real-time tracking at a high frame rate – track mode is cheaper than register, but this is still a deep-learning pipeline running per frame, not a lightweight visual tracker
  • The object is heavily occluded or its bbox can't be tightly drawn – registration accuracy depends on a reasonably clean, tight bbox

TIP

If tracking starts drifting after many frames, re-register with a fresh bbox rather than continuing to feed in a stale tracking_pose - small per-frame errors can accumulate.