Metric depth
Dense image-aligned depth recovers scene surfaces in absolute scale across different camera configurations.
We present a Geometry-grounded Unified 3D Perception (GeoUP) framework, which adapts the reconstruction-oriented latent of VGGT to calibrated, streaming multi-camera driving scenes and uses one geometry-grounded representation for metric depth estimation, 3D object detection, and semantic occupancy prediction.
Dense image-aligned depth recovers scene surfaces in absolute scale across different camera configurations.
Object queries read metric location, extent, orientation, and velocity from the shared geometric latent.
Sparse occupancy queries recover voxel-level scene state and semantics from multi-frame, multi-view features.
GeoUP combines image patch tokens with calibration-derived raymap embeddings and camera tokens. A shared Transformer alternates self, temporal, and view attention before task-specific heads decode depth, detection, occupancy, and auxiliary camera pose predictions.
Patch-aligned Plücker rays derived from intrinsics and poses inject camera geometry and absolute metric scale.
Temporal attention follows each camera stream while view attention exchanges information across synchronized cameras.
DPT-, RayDN-, and OPUS-V2-style heads read different geometric levels without fragmenting the shared backbone.
GeoUP first adapts geometry-pretrained VGGT features to driving scenes with depth and camera supervision. It then jointly learns detection, occupancy, depth, and camera objectives across five datasets. Task masks allow every sample to supervise only the annotations it provides.
GeoUP achieves state-of-the-art results across 3D detection, semantic occupancy, and metric depth estimation. Joint multi-dataset training further improves detection on nuScenes, Argoverse 2, and Waymo, while also strengthening occupancy quality on Occ3D-nuScenes. GeoUP (Joint) denotes the model trained jointly on nuScenes, Argoverse 2, Waymo, DDAD, and KITTI.
Compare instance-level detection, volume-level semantic occupancy, and surface-level depth reconstruction.
Please cite our paper if GeoUP is useful for your research.
@article{xu2026geometry,
title={Geometry-Grounded Unified 3D Perception for Autonomous Driving},
author={Xu, Longfei and Wang, Xiaohui and Huang, Zehao and Li, Han and Yang, Ya and Wang, Naiyan and Liu, Si},
journal={arXiv preprint arXiv:2608.13147},
year={2026}
}