Skip to content

Estimate Depth Using Fast-FoundationStereo ​

SUMMARY

Estimate Depth Using Fast-FoundationStereo predicts a metric depth map from a rectified left/right stereo image pair.

It runs NVIDIA's Fast-FoundationStereo, a real-time, zero-shot stereo-matching model, over the two images plus the left camera's intrinsics and the stereo baseline, and returns a datatypes.DepthImage. Unlike classical block-matching stereo, it needs no per-scene tuning and generalizes to unseen scenes out of the box. Chain the result into convert_depth_image_to_point_cloud to get a 3D point cloud.

Use this Skill when you have a calibrated, rectified stereo camera pair and want a metric depth map without a dedicated depth sensor.

Example ​

Left Stereo Image

Fast-FoundationStereo left image

The left image from the rectified stereo pair.

Right Stereo Image

Fast-FoundationStereo right image

The right image used with the left image to estimate depth.

Estimated Depth

Fast-FoundationStereo depth image

The metric depth map predicted from the rectified stereo pair.

The Skill ​

python
from telekinesis import vitreous

depth_image = vitreous.estimate_depth_using_fast_foundation_stereo(
    left_image=left_image,
    right_image=right_image,
    camera_calibration=camera_calibration,
    baseline=baseline,
)
API Reference
Full parameter and return type documentation for estimate_depth_using_fast_foundation_stereo.
View Reference →

The Code ​

python
"""
Demonstrates estimating a metric depth map from a rectified stereo image
pair using Fast-FoundationStereo.
"""

import numpy as np
from loguru import logger
import rerun as rr

from telekinesis import vitreous, datatypes


def estimate_depth_using_fast_foundation_stereo_example():
    """
    Estimates a metric depth map from a synthetic rectified stereo pair.

    Builds a left image with a random texture and a right image that is
    the same texture shifted by a fixed pixel disparity, simulating a
    scene at a single known depth (with this camera/baseline, a 20px
    disparity corresponds to 1.5m).
    """
    # ===================== Load Data ==========================================
    height, width = 480, 640
    disparity_px = 20
    baseline = 0.05  # meters

    rng = np.random.default_rng(0)
    texture = rng.integers(
        0, 255, (height, width + disparity_px, 3), dtype=np.uint8
    )
    left_image = datatypes.Image(texture[:, :width])
    right_image = datatypes.Image(texture[:, disparity_px : disparity_px + width])

    camera_calibration = datatypes.CameraCalibration(
        width=width,
        height=height,
        distortion_model="plumb_bob",
        distortion_parameters=[0.0, 0.0, 0.0, 0.0, 0.0],
        intrinsic_matrix=[600.0, 0.0, 320.0, 0.0, 600.0, 240.0, 0.0, 0.0, 1.0],
    )

    # ===================== Run Skill ==========================================
    depth_image = vitreous.estimate_depth_using_fast_foundation_stereo(
        left_image=left_image,
        right_image=right_image,
        camera_calibration=camera_calibration,
        baseline=baseline,
    )

    # ===================== Log ================================================
    logger.success(f"Estimated depth from {left_image} and {right_image}")
    logger.success(f"Results: {depth_image}")
    logger.info(f"Depth image shape: {depth_image.shape}")

    # ===================== Visualization  (Optional) ===========================
    rr.init("estimate_depth_using_fast_foundation_stereo_example", spawn=True)
    datatypes.visualize(left_image, entity_path="/1-left_image")
    datatypes.visualize(right_image, entity_path="/2-right_image")
    datatypes.visualize(depth_image, entity_path="/3-depth_image")


if __name__ == "__main__":
    estimate_depth_using_fast_foundation_stereo_example()

Runnable examples are available in the Telekinesis examples repository.

Follow the README in that repository to set up the environment, run this specific example with:

bash
cd telekinesis-examples
python examples/point_cloud/estimate_depth_using_fast_foundation_stereo.py

Parameter Configuration ​

ParameterTypeDefaultDescription
left_imagedatatypes.Image | np.ndarrayrequiredThe left camera's image, shape (H, W, 3).
right_imagedatatypes.Image | np.ndarrayrequiredThe right camera's image, already rectified to left_image (same resolution, row-aligned epipolar lines).
camera_calibrationdatatypes.CameraCalibration | dictrequiredThe left camera's intrinsic calibration. Only intrinsic_matrix is used by this Skill.
baselinedatatypes.Float | floatrequiredThe stereo baseline between the left and right cameras, in meters.
valid_itersdatatypes.Int | int8Number of GRU refinement iterations. This model was distilled for real-time use at its default; increasing it further is not guaranteed to improve accuracy the way it would for a non-distilled stereo model.
max_dispdatatypes.Int | int192Maximum disparity search range, in pixels. Increase for scenes with larger disparities (e.g. very close objects with a wide baseline).

Returns ​

TypeDescription
datatypes.DepthImageA metric depth map, shape (H, W), in meters, with left_image attached as its aligned color plane. Zero marks pixels with no valid depth. Use .depth for the raw depth array and .colors for the aligned color plane.

Raises ​

ExceptionCondition
TypeErrorA parameter's value does not match its expected type (see the Parameter Configuration table above)
ValueErrorleft_image and right_image don't share the same size, or baseline is not positive
ConfigurationErrorThe TELEKINESIS_API_KEY environment variable is not set
SerializationErrorThe request input failed to serialize, the response was not returned as an Arrow stream, or the response failed to deserialize
RequestTimeoutErrorThe request to the Vitreous service timed out
TransportErrorA network failure occurred before a response was received
ClientErrorThe Vitreous service rejected the request due to invalid or malformed input (HTTP 400/422), an unrecognized endpoint (HTTP 404), or another unexpected 4xx response
AuthenticationErrorThe API key was rejected as invalid or expired (HTTP 401)
AuthenticationServiceErrorThe authentication service returned an invalid response, was temporarily unavailable, or timed out (HTTP 502/503/504)
ServerErrorThe Vitreous service returned a 5xx or otherwise unexpected error response

How to Tune the Parameters ​

estimate_depth_using_fast_foundation_stereo has two knobs beyond the required inputs:

  • valid_iters controls how many GRU refinement iterations run. This is a distilled, real-time stereo model tuned specifically to converge in its default 8 iterations — pushing it higher is not guaranteed to help the way it would for a full-size stereo network, and mainly costs compute time.
  • max_disp bounds the disparity search range. If your scene has objects close enough (relative to the baseline) that their true disparity exceeds this value, depth for those pixels will be wrong or missing — increase it for close-range / wide-baseline setups.

Beyond those, accuracy depends entirely on how well-calibrated and well-rectified your stereo pair is:

  • camera_calibration and baseline must be accurate. Depth is derived from disparity via fx * baseline / disparity, so an incorrect focal length or baseline directly scales the resulting depth.
  • left_image/right_image must already be rectified (row-aligned epipolar lines). This Skill does not perform rectification itself.

Where to Use the Skill ​

Common pipelines include:

  • Stereo-camera perception – getting a metric depth map from a calibrated stereo rig without a dedicated depth sensor
  • RGB-D pipeline replacement – producing depth for scenes/materials (e.g. transparent or reflective objects) where active depth sensors (structured light, ToF) struggle but passive stereo still works
  • Point cloud generation – chaining the result into convert_depth_image_to_point_cloud for downstream point-cloud processing, registration, or pose estimation

Alternative Skills ​

Skillvs. Estimate Depth Using Fast-FoundationStereo
convert_depth_image_to_point_cloudTakes a DepthImage you already have (from any source) and back-projects it into a PointCloud. Use it right after this Skill to get 3D points instead of a depth map.

When Not to Use the Skill ​

Do not use Estimate Depth Using Fast-FoundationStereo when:

  • You already have a depth sensor (structured light, ToF, or an active stereo camera that outputs depth directly) – there's no need to run stereo matching yourself
  • Your stereo pair is not rectified – this Skill assumes row-aligned epipolar lines; rectify the images first
  • You don't know your camera's intrinsics or baseline – depth accuracy depends directly on both; without a real calibration, the output will be scaled incorrectly

TIP

If the resulting depth looks scaled incorrectly (too close/too far by a consistent factor), double-check baseline and camera_calibration.intrinsic_matrix before assuming the Skill itself is wrong — depth is directly proportional to both.