ExStereo Lifting 2D Vision-Language-Action Models to 3D with Explicit Stereo Representations

Anonymous Authors

Overview Video

Abstract

Stereo perception for off-the-shelf VLAs

Three-dimensional perception is critical for robotic manipulation, particularly for high-precision tasks, as recovering metric depth and precise 3D object positions from monocular RGB observations is inherently ill-posed. However, many Vision-Language-Action (VLA) models rely solely on RGB observations for perception.

Leveraging recent advances in foundation models for stereo matching, we introduce ExStereo, a stereo module that augments pre-trained 2D VLAs with 3D perception. ExStereo reconstructs scene geometry from stereo image pairs and renders multi-view observations as an explicit stereo representation for stereo feature extraction. The action tokens from the action expert selectively attend to the resulting stereo tokens through our proposed action-stereo cross-attention mechanism, enabling the policy to generate robot actions conditioned on 3D scene information. To learn robust 3D representations, we introduce a mid-training stage before task-specific post-training, using a self-supervised learning objective on large-scale stereo data.

We validate our approach by fine-tuning two publicly available VLAs, π0.5 and SmolVLA, and evaluate them in simulation and on a real-world bimanual PiPER platform. Across both settings, VLAs fine-tuned with ExStereo consistently outperform baselines, demonstrating the effectiveness of stereo perception for robotic manipulation.

A stereo camera's left and right images feed ExStereo, which builds a four-view explicit stereo representation (Top, Front, Left, Right) that the action expert of an off-the-shelf VLA reads through action-stereo cross-attention.
Given a stereo pair, ExStereo constructs four orthogonal views of the scene with metric 3D coordinates. This explicit 3D representation allows the VLAs to selectively retrieve visual features relevant to the action tokens through our proposed action-stereo cross-attention mechanism, improving robotic manipulation performance.

Method Overview

From a stereo pair to 3D-conditioned actions

ExStereo is a stereo module that equips pre-trained, off-the-shelf VLAs with 3D spatial perception. It takes a stereo image pair as input and employs a pre-trained stereo foundation model to obtain a disparity map. Using the camera intrinsics and stereo baseline, the disparity map is transformed into a metric depth map and 3D point cloud. The point cloud is rendered from four orthogonal viewpoints (Top, Front, Left, and Right) using z-buffering. The four views are then tiled horizontally to form a multi-view observation, which is fed into a Vision Transformer (ViT) to obtain stereo tokens. The stereo tokens are utilized for action generation through the proposed action-stereo cross-attention mechanism, where the action tokens serve as queries and the stereo tokens serve as keys and values. This allows the action tokens to selectively attend to relevant 3D visual features without modifying the large prefix token sequence or the weights of VLMs, enabling pre-trained robot policies to efficiently integrate the stereo module during post-training.

ExStereo pipeline: the stereo foundation model predicts disparity, which becomes a 3D point cloud rendered into four views; a Vision Transformer produces stereo tokens that action tokens query through cross-attention. During mid-training, a masked-autoencoding decoder reconstructs masked views.
Overview of ExStereo
1

Explicit stereo representation

Stereo foundation model predicts a disparity map, which is converted to metric depth and back-projected into a point cloud. After RANSAC table leveling, it renders four 128×128 orthographic views (Top, Front, Left, Right) with depth, XYZ, and validity, tiled into a 128×512×5 observation.

2

Action-stereo cross-attention

A ViT with 16×16 patches turns the observation into 256 unpooled stereo tokens. A single multi-head cross-attention block lets each action token query them. Its output is added back to the action tokens as a residual, so the pre-trained VLM is untouched.

3

Mid-training, then post-training

During mid-training, the VLA with ExStereo train on large-scale stereo data with the standard action loss plus a masked-autoencoding (MAE) loss on the multi-view patches. During post-training, the stereo module is frozen, the MAE decoder is dropped, and the policy is fine-tuned with LoRA on a small set of task demos.

Transparent glass bottles: an RGB image, a noisy RealSense depth point cloud with missing geometry, and a clean point cloud reconstructed from a ZED Mini stereo pair.

Why stereo?

Geometry where depth sensors fail

RGB-D sensors struggle with transparent and reflective objects. On these glass bottles, a RealSense point cloud (bottom left) is noisy and full of holes. A ZED Mini stereo pair processed with Fast-FoundationStereo (bottom right) recovers clean, consistent geometry, with no extra depth hardware.

Simulation

Simulation: RoboTwin bimanual tasks

We fine-tune π0.5 and SmolVLA with ExStereo and compare them with their fine-tuned base policies, with ACT (RGB), DP3 (point cloud) and StereoVLA (implicit stereo), and with ExStereo without mid-training. Each method is evaluated on 25 unseen, domain-randomized episodes, and results are averaged over three training seeds.

Ten simulated RoboTwin tasks with domain randomization.
Top: Lift Pot, Adjust Bottle, Grab Roller, Open Laptop, Beat Block Hammer. Bottom: Click Alarmclock, Dump Bin Bigbin, Handover Mic, Press Stapler, Stamp Seal.
π0.5 avg. success
51.2 → 76.4
+25.2 points with ExStereo
SmolVLA avg. success
18.7 → 42.0
+23.3 points with ExStereo
Largest single-task gain
+65.3
SmolVLA on Handover Mic (0.0 → 65.3)
Success rate (%) on the selected task. Blue bars are VLAs fine-tuned with ExStereo.

π0.5-ExStereo rollouts

Evaluation episodes on unseen, domain-randomized scenes. Pick a task to watch its rollouts.

Real World

Bimanual AgileX PiPER with a ZED Mini

We test on three real-world tasks: a standard bimanual task (Stack Two Bowls), a task with transparent objects (Lift Glass Bottles), and a task with an articulated object (Close Fridge Door). Each method and task gets 20 trials.

Successful trials on the selected task (20 trials per task). Blue bars are π0.5 fine-tuned with ExStereo. “w/o MT” denotes ExStereo without the mid-training stage.

Real-world videos