Estimate Pose Using FoundationPose
SUMMARY
Estimate Pose Using FoundationPose estimates a 6-DOF object pose from an RGB-D frame plus a CAD mesh of the object.
It runs NVIDIA's FoundationPose, a zero-shot pose estimation and tracking model, in one of two modes: register (pass a bbox) samples and scores many pose hypotheses to estimate a pose from scratch; track (pass previous_pose) refines a prior estimate for a new frame, much more cheaply. To track an object across frames, register once, then feed each call's tracking_pose back in as the next call's previous_pose.
Use this Skill when you have a CAD mesh of an object and want its 6-DOF pose for grasping, manipulation, or scene understanding, without training a per-object detector.
The Skill
from telekinesis import vitreous
# Register on the first frame.
pose, tracking_pose = vitreous.estimate_pose_using_foundation_pose(
rgb_image=rgb_image,
depth_image=depth_image,
mesh=mesh,
camera_calibration=camera_calibration,
bbox=bbox,
)
# Track on later frames.
pose, tracking_pose = vitreous.estimate_pose_using_foundation_pose(
rgb_image=next_rgb_image,
depth_image=next_depth_image,
mesh=mesh,
camera_calibration=camera_calibration,
previous_pose=tracking_pose,
)The Code
"""
Demonstrates estimating a 6-DOF object pose from an RGB-D frame and a CAD
mesh using FoundationPose, then tracking it across a second frame.
"""
import numpy as np
from loguru import logger
import rerun as rr
from telekinesis import vitreous, datatypes
def estimate_pose_using_foundation_pose_example():
"""
Registers a pose on a synthetic first frame, then tracks it on a
second frame using the returned tracking_pose.
"""
# ===================== Load Data ==========================================
height, width = 480, 640
rgb_image = np.full((height, width, 3), 200, dtype=np.uint8)
depth_image = np.full((height, width), 1.0, dtype=np.float32) # a flat wall 1m away
# A small cube mesh, centered at the origin.
mesh = datatypes.Mesh3D(
vertex_positions=[
[-0.05, -0.05, -0.05], [0.05, -0.05, -0.05], [0.05, 0.05, -0.05], [-0.05, 0.05, -0.05],
[-0.05, -0.05, 0.05], [0.05, -0.05, 0.05], [0.05, 0.05, 0.05], [-0.05, 0.05, 0.05],
],
triangle_indices=[
[0, 1, 2], [0, 2, 3], [4, 6, 5], [4, 7, 6],
[0, 4, 5], [0, 5, 1], [1, 5, 6], [1, 6, 2],
[2, 6, 7], [2, 7, 3], [3, 7, 4], [3, 4, 0],
],
)
camera_calibration = datatypes.CameraCalibration(
width=width,
height=height,
distortion_model="plumb_bob",
distortion_parameters=[0.0, 0.0, 0.0, 0.0, 0.0],
intrinsic_matrix=[600.0, 0.0, 320.0, 0.0, 600.0, 240.0, 0.0, 0.0, 1.0],
)
bbox = [280, 200, 360, 280] # a small region roughly at the image center
# ===================== Run Skill ==========================================
# Frame 1: register from scratch using the bbox.
pose, tracking_pose = vitreous.estimate_pose_using_foundation_pose(
rgb_image=rgb_image,
depth_image=depth_image,
mesh=mesh,
camera_calibration=camera_calibration,
bbox=bbox,
)
# Frame 2: track forward using the previous frame's tracking_pose.
pose, tracking_pose = vitreous.estimate_pose_using_foundation_pose(
rgb_image=rgb_image,
depth_image=depth_image,
mesh=mesh,
camera_calibration=camera_calibration,
previous_pose=tracking_pose,
)
# ===================== Log ================================================
logger.success(f"Estimated pose for {mesh}")
logger.success(f"Results: {pose}")
logger.info(f"Pose matrix:\n{pose.data}")
# ===================== Visualization (Optional) ===========================
rr.init("estimate_pose_using_foundation_pose_example", spawn=True)
datatypes.visualize(mesh, entity_path="/1-mesh")
if __name__ == "__main__":
estimate_pose_using_foundation_pose_example()Runnable examples are available in the Telekinesis examples repository.
Follow the README in that repository to set up the environment, run this specific example with:
cd telekinesis-examples
python examples/point_cloud/estimate_pose_using_foundation_pose.pyParameter Configuration
| Parameter | Type | Default | Description |
|---|---|---|---|
rgb_image | datatypes.Image | np.ndarray | required | The RGB image, shape (H, W, 3). |
depth_image | datatypes.DepthImage | np.ndarray | required | The metric depth map, aligned to rgb_image, shape (H, W), in meters. |
mesh | datatypes.Mesh3D | dict | required | The tracked object's CAD mesh. Only vertex_positions/triangle_indices are used. |
camera_calibration | datatypes.CameraCalibration | dict | required | The camera's intrinsic calibration. Only intrinsic_matrix is used. |
bbox | datatypes.Box2D | np.ndarray | list | None | 2D bounding box [umin, vmin, umax, vmax] (XYXY) around the object, used to register a new pose. Exactly one of bbox/previous_pose must be given. |
previous_pose | datatypes.Transform3D | np.ndarray | None | A prior call's tracking_pose, used to track an existing pose. Exactly one of bbox/previous_pose must be given. |
refine_iterations | datatypes.Int | int | 2 | Number of refine iterations to run. |
Returns
| Type | Description |
|---|---|
(pose, tracking_pose) tuple of two datatypes.Transform3D | pose is the object-to-camera transform in the mesh's own (uncentered) coordinate system - the pose your application should use. tracking_pose is opaque continuation state; pass it back as previous_pose on a later call to track this object more cheaply than re-registering. Don't interpret its values directly. |
Raises
| Exception | Condition |
|---|---|
TypeError | A parameter's value does not match its expected type (see the Parameter Configuration table above) |
ValueError | Neither or both of bbox/previous_pose are given |
ConfigurationError | The TELEKINESIS_API_KEY environment variable is not set |
SerializationError | The request input failed to serialize, the response was not returned as an Arrow stream, or the response failed to deserialize |
RequestTimeoutError | The request to the Vitreous service timed out |
TransportError | A network failure occurred before a response was received |
ClientError | The Vitreous service rejected the request due to invalid or malformed input (HTTP 400/422), an unrecognized endpoint (HTTP 404), or another unexpected 4xx response |
AuthenticationError | The API key was rejected as invalid or expired (HTTP 401) |
AuthenticationServiceError | The authentication service returned an invalid response, was temporarily unavailable, or timed out (HTTP 502/503/504) |
ServerError | The Vitreous service returned a 5xx or otherwise unexpected error response |
How to Tune the Parameters
bboxvs.previous_poseis the main mode switch: usebboxon the first frame (or whenever you've lost track of the object), andprevious_poseon subsequent frames of the same object to track it much more cheaply than re-registering.refine_iterationscontrols how many refinement passes run per call, for both register and track. More iterations can marginally improve pose accuracy at the cost of compute time.bboxshould tightly enclose the object when registering - a loose or offset box widens the search and can lower accuracy or cause registration to fail.- Track only holds up if the object hasn't moved far between frames relative to your capture rate; if tracking drifts or loses the object, re-register with a fresh
bbox.
Where to Use the Skill
Common pipelines include:
- Robotic grasping – getting an object's 6-DOF pose to plan a grasp
- Bin picking – registering pose for an object identified by a detector, then tracking it through a pick-and-place motion
- Scene understanding – localizing known CAD objects in a scene for downstream planning
Alternative Skills
| Skill | vs. Estimate Pose Using FoundationPose |
|---|---|
| estimate_depth_using_fast_foundation_stereo | Produces a metric depth map from a stereo pair, rather than an object pose. Useful as an input source for depth_image if you don't already have a depth sensor. |
When Not to Use the Skill
Do not use Estimate Pose Using FoundationPose when:
- You don't have a CAD mesh of the object – this Skill requires a mesh; if you only need to detect/segment the object (not its 3D pose), use a Cornea or Retina Skill instead
- You need real-time tracking at a high frame rate – track mode is cheaper than register, but this is still a deep-learning pipeline running per frame, not a lightweight visual tracker
- The object is heavily occluded or its bbox can't be tightly drawn – registration accuracy depends on a reasonably clean, tight
bbox
TIP
If tracking starts drifting after many frames, re-register with a fresh bbox rather than continuing to feed in a stale tracking_pose - small per-frame errors can accumulate.