Geometry-Grounded Unified 3D Perception for Autonomous Driving

Geometry-Grounded Unified 3D Perception for Autonomous Driving

Longfei Xu1,* Xiaohui Wang2,* Zehao Huang Han Li1 Ya Yang2 Naiyan Wang Si Liu1

* equal contribution

1. Beihang University 2. Beijing University of Posts and Telecommunications

We present a Geometry-grounded Unified 3D Perception (GeoUP) framework, which adapts the reconstruction-oriented latent of VGGT to calibrated, streaming multi-camera driving scenes and uses one geometry-grounded representation for metric depth estimation, 3D object detection, and semantic occupancy prediction.

Why reconstruction pretraining? It provides semantic features, explicit geometry, and multi-image consistency in one representation.
Surface level

Metric depth

Dense image-aligned depth recovers scene surfaces in absolute scale across different camera configurations.

Instance level

3D object detection

Object queries read metric location, extent, orientation, and velocity from the shared geometric latent.

Volume level

Semantic occupancy

Sparse occupancy queries recover voxel-level scene state and semantics from multi-frame, multi-view features.

The Model

GeoUP combines image patch tokens with calibration-derived raymap embeddings and camera tokens. A shared Transformer alternates self, temporal, and view attention before task-specific heads decode depth, detection, occupancy, and auxiliary camera pose predictions.

Overall GeoUP pipeline from streaming multi-view images to unified 3D perception outputs
Overall pipeline. Geometry-aware tokens are processed by a factorized spatiotemporal backbone and decoded into heterogeneous 3D perception outputs.
Geometry injection

Calibration-aware raymaps

Patch-aligned Plücker rays derived from intrinsics and poses inject camera geometry and absolute metric scale.

Structured interaction

Temporal-view factorization

Temporal attention follows each camera stream while view attention exchanges information across synchronized cameras.

Unified readout

Multi-task perception heads

DPT-, RayDN-, and OPUS-V2-style heads read different geometric levels without fragmenting the shared backbone.

Unified Training

GeoUP first adapts geometry-pretrained VGGT features to driving scenes with depth and camera supervision. It then jointly learns detection, occupancy, depth, and camera objectives across five datasets. Task masks allow every sample to supervise only the annotations it provides.

nuScenes Argoverse 2 Waymo DDAD KITTI
Dataset sampling

Joint training ratio

Training strategy

From geometry prior to unified perception

01 VGGT Initialization
02 Driving Adaptation
03 Datasets Align
04 Multi-task Learning
05 Geometry-grounded Latent

Unified 3D Perception Results

GeoUP achieves state-of-the-art results across 3D detection, semantic occupancy, and metric depth estimation. Joint multi-dataset training further improves detection on nuScenes, Argoverse 2, and Waymo, while also strengthening occupancy quality on Occ3D-nuScenes. GeoUP (Joint) denotes the model trained jointly on nuScenes, Argoverse 2, Waymo, DDAD, and KITTI.

Five driving benchmarks

Detection, occupancy, and depth

Citation

Please cite our paper if GeoUP is useful for your research.

@article{xu2026geometry,
  title={Geometry-Grounded Unified 3D Perception for Autonomous Driving},
  author={Xu, Longfei and Wang, Xiaohui and Huang, Zehao and Li, Han and Yang, Ya and Wang, Naiyan and Liu, Si},
  journal={arXiv preprint arXiv:2608.13147},
  year={2026}
}