用点云预测未来场景结构,提升自动驾驶决策能力。
GeoWAM: Visual Geometry World Action Models for Autonomous Driving

- 以点云代替图像建模场景演化,更直接反映三维空间关系
- 在多个评测中优于基于图像的模型,显著提升驾驶策略性能
- 适合研究自动驾驶感知与规划融合的开发者参考
世界动作模型(WAMs)近年来受到关注,用于联合建模自动驾驶中的场景演化与车辆自身动作。现有方法大多在像素空间学习场景动态,结合视频生成主干网络预测未来观测,并通过动作头预测车辆轨迹。然而,像素仅间接表示动态:它将几何、运动与外观、纹理、光照混杂在一起,迫使模型从二维观察推断三维变换。我们认为,以点云表示的几何结构是更自然的状态空间,因其显式捕捉空间结构及刚性与非刚性变换,且与驾驶动作执行空间直接对齐。基于此,我们提出【GeoWAM】——一种面向自动驾驶的视觉几何世界动作模型。不同于预测未来图像,GeoWAM 预训练为预测未来场景几何,生成同时编码空间结构与时间演化的表征。随后,一个几何条件的动作头利用这些学习到的几何动态,预测未来的车辆轨迹。大量开环与闭环评估表明,视觉几何世界建模显著优于基于图像的替代方案,确立未来几何预测作为自动驾驶有效预训练目标。
原文摘要 · Abstract (English)
World action models (WAMs) have recently gained increasing attention as a framework for jointly modeling scene evolution and ego actions in autonomous driving. Most existing WAMs learn scene dynamics in pixel space by combining a video-generation backbone for future-observation prediction with an action head for ego-trajectory prediction. Pixels, however, provide only an indirect representation of these dynamics: they entangle geometry and motion with appearance, texture, and illumination, forcing the model to infer three-dimensional transformations from two-dimensional observations. We argue that geometry, represented by point clouds, offers a more natural state space for driving because it explicitly captures spatial structure and the rigid and non-rigid transformations that govern scene evolution while directly aligning with the space in which driving actions are executed. Building on this insight, we introduce \textbf{GeoWAM}, a visual geometry world action model for autonomous driving. Rather than predicting future images, GeoWAM is pretrained to forecast future scene geometry, yielding representations that jointly encode spatial structure and temporal evolution. A geometry-conditioned action head then leverages these learned geometric dynamics to predict future ego trajectories. Extensive open-loop and closed-loop evaluations show that visual geometry world modeling yields substantially stronger driving policies than image-based alternatives, establishing future-geometry prediction as an effective pretraining objective for autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。