Estimate Depth Using Fast-FoundationStereo
SUMMARY
Estimate Depth Using Fast-FoundationStereo predicts a metric depth map from a rectified left/right stereo image pair.
It runs NVIDIA's Fast-FoundationStereo, a real-time, zero-shot stereo-matching model, over the two images plus the left camera's intrinsics and the stereo baseline, and returns a datatypes.DepthImage. Unlike classical block-matching stereo, it needs no per-scene tuning and generalizes to unseen scenes out of the box. Chain the result into convert_depth_image_to_point_cloud to get a 3D point cloud.
Use this Skill when you have a calibrated, rectified stereo camera pair and want a metric depth map without a dedicated depth sensor.
Example
Left Stereo Image

The left image from the rectified stereo pair.
Right Stereo Image

The right image used with the left image to estimate depth.
Estimated Depth

The metric depth map predicted from the rectified stereo pair.
The Skill
from telekinesis import vitreous
depth_image = vitreous.estimate_depth_using_fast_foundation_stereo(
left_image=left_image,
right_image=right_image,
camera_calibration=camera_calibration,
baseline=baseline,
)The Code
"""
Demonstrates estimating a metric depth map from a rectified stereo image
pair using Fast-FoundationStereo.
"""
import numpy as np
from loguru import logger
import rerun as rr
from telekinesis import vitreous, datatypes
def estimate_depth_using_fast_foundation_stereo_example():
"""
Estimates a metric depth map from a synthetic rectified stereo pair.
Builds a left image with a random texture and a right image that is
the same texture shifted by a fixed pixel disparity, simulating a
scene at a single known depth (with this camera/baseline, a 20px
disparity corresponds to 1.5m).
"""
# ===================== Load Data ==========================================
height, width = 480, 640
disparity_px = 20
baseline = 0.05 # meters
rng = np.random.default_rng(0)
texture = rng.integers(
0, 255, (height, width + disparity_px, 3), dtype=np.uint8
)
left_image = datatypes.Image(texture[:, :width])
right_image = datatypes.Image(texture[:, disparity_px : disparity_px + width])
camera_calibration = datatypes.CameraCalibration(
width=width,
height=height,
distortion_model="plumb_bob",
distortion_parameters=[0.0, 0.0, 0.0, 0.0, 0.0],
intrinsic_matrix=[600.0, 0.0, 320.0, 0.0, 600.0, 240.0, 0.0, 0.0, 1.0],
)
# ===================== Run Skill ==========================================
depth_image = vitreous.estimate_depth_using_fast_foundation_stereo(
left_image=left_image,
right_image=right_image,
camera_calibration=camera_calibration,
baseline=baseline,
)
# ===================== Log ================================================
logger.success(f"Estimated depth from {left_image} and {right_image}")
logger.success(f"Results: {depth_image}")
logger.info(f"Depth image shape: {depth_image.shape}")
# ===================== Visualization (Optional) ===========================
rr.init("estimate_depth_using_fast_foundation_stereo_example", spawn=True)
datatypes.visualize(left_image, entity_path="/1-left_image")
datatypes.visualize(right_image, entity_path="/2-right_image")
datatypes.visualize(depth_image, entity_path="/3-depth_image")
if __name__ == "__main__":
estimate_depth_using_fast_foundation_stereo_example()Runnable examples are available in the Telekinesis examples repository.
Follow the README in that repository to set up the environment, run this specific example with:
cd telekinesis-examples
python examples/point_cloud/estimate_depth_using_fast_foundation_stereo.pyParameter Configuration
| Parameter | Type | Default | Description |
|---|---|---|---|
left_image | datatypes.Image | np.ndarray | required | The left camera's image, shape (H, W, 3). |
right_image | datatypes.Image | np.ndarray | required | The right camera's image, already rectified to left_image (same resolution, row-aligned epipolar lines). |
camera_calibration | datatypes.CameraCalibration | dict | required | The left camera's intrinsic calibration. Only intrinsic_matrix is used by this Skill. |
baseline | datatypes.Float | float | required | The stereo baseline between the left and right cameras, in meters. |
valid_iters | datatypes.Int | int | 8 | Number of GRU refinement iterations. This model was distilled for real-time use at its default; increasing it further is not guaranteed to improve accuracy the way it would for a non-distilled stereo model. |
max_disp | datatypes.Int | int | 192 | Maximum disparity search range, in pixels. Increase for scenes with larger disparities (e.g. very close objects with a wide baseline). |
Returns
| Type | Description |
|---|---|
datatypes.DepthImage | A metric depth map, shape (H, W), in meters, with left_image attached as its aligned color plane. Zero marks pixels with no valid depth. Use .depth for the raw depth array and .colors for the aligned color plane. |
Raises
| Exception | Condition |
|---|---|
TypeError | A parameter's value does not match its expected type (see the Parameter Configuration table above) |
ValueError | left_image and right_image don't share the same size, or baseline is not positive |
ConfigurationError | The TELEKINESIS_API_KEY environment variable is not set |
SerializationError | The request input failed to serialize, the response was not returned as an Arrow stream, or the response failed to deserialize |
RequestTimeoutError | The request to the Vitreous service timed out |
TransportError | A network failure occurred before a response was received |
ClientError | The Vitreous service rejected the request due to invalid or malformed input (HTTP 400/422), an unrecognized endpoint (HTTP 404), or another unexpected 4xx response |
AuthenticationError | The API key was rejected as invalid or expired (HTTP 401) |
AuthenticationServiceError | The authentication service returned an invalid response, was temporarily unavailable, or timed out (HTTP 502/503/504) |
ServerError | The Vitreous service returned a 5xx or otherwise unexpected error response |
How to Tune the Parameters
estimate_depth_using_fast_foundation_stereo has two knobs beyond the required inputs:
valid_iterscontrols how many GRU refinement iterations run. This is a distilled, real-time stereo model tuned specifically to converge in its default 8 iterations — pushing it higher is not guaranteed to help the way it would for a full-size stereo network, and mainly costs compute time.max_dispbounds the disparity search range. If your scene has objects close enough (relative to the baseline) that their true disparity exceeds this value, depth for those pixels will be wrong or missing — increase it for close-range / wide-baseline setups.
Beyond those, accuracy depends entirely on how well-calibrated and well-rectified your stereo pair is:
camera_calibrationandbaselinemust be accurate. Depth is derived from disparity viafx * baseline / disparity, so an incorrect focal length or baseline directly scales the resulting depth.left_image/right_imagemust already be rectified (row-aligned epipolar lines). This Skill does not perform rectification itself.
Where to Use the Skill
Common pipelines include:
- Stereo-camera perception – getting a metric depth map from a calibrated stereo rig without a dedicated depth sensor
- RGB-D pipeline replacement – producing depth for scenes/materials (e.g. transparent or reflective objects) where active depth sensors (structured light, ToF) struggle but passive stereo still works
- Point cloud generation – chaining the result into
convert_depth_image_to_point_cloudfor downstream point-cloud processing, registration, or pose estimation
Alternative Skills
| Skill | vs. Estimate Depth Using Fast-FoundationStereo |
|---|---|
| convert_depth_image_to_point_cloud | Takes a DepthImage you already have (from any source) and back-projects it into a PointCloud. Use it right after this Skill to get 3D points instead of a depth map. |
When Not to Use the Skill
Do not use Estimate Depth Using Fast-FoundationStereo when:
- You already have a depth sensor (structured light, ToF, or an active stereo camera that outputs depth directly) – there's no need to run stereo matching yourself
- Your stereo pair is not rectified – this Skill assumes row-aligned epipolar lines; rectify the images first
- You don't know your camera's intrinsics or baseline – depth accuracy depends directly on both; without a real calibration, the output will be scaled incorrectly
TIP
If the resulting depth looks scaled incorrectly (too close/too far by a consistent factor), double-check baseline and camera_calibration.intrinsic_matrix before assuming the Skill itself is wrong — depth is directly proportional to both.