轻量级3D场景重建模型,实现实时自动驾驶环境感知。
LiAuto-GeoX: Efficient Grounded Driving Transformer

- 用稀疏激光雷达引导,实现远距离几何精准建模。
- 155M参数模型达220帧/秒,保持高精度三维重建。
- 适用于自动驾驶的实时感知与下游任务部署。
稠密3D重建在空间理解中潜力巨大,但其在动态驾驶环境中作为实时、车载表示的可行性仍面临挑战。现有大规模视觉几何模型通常需要大量计算资源,且缺乏长程几何保真度、环视一致性与实时效率。为此,我们提出LiAuto-GeoX,一种高效、基于地面的驾驶变换器,用于可部署的自车中心3D场景理解。方法从大规模环视数据中学习高容量驾驶几何模型,并利用稀疏激光雷达先验,在远距离、模糊或结构稀疏区域提供稳健的几何锚定。通过新颖的几何保真蒸馏框架,将该能力压缩为仅155M参数的车载模型。该框架采用掩码引导的深度感知蒸馏以保留细粒度度量结构,强调几何信息丰富的区域;以及相对位姿关系蒸馏,通过位姿诱导的几何关系强化跨视角空间一致性。大量实验表明,LiAuto-GeoX在KITTI上达到220 FPS,同时保持高保真稠密重建,支持实时部署。所学几何可无缝迁移至下游自主任务:轨迹预测达90.6 PDMS,占用预测达24.63 mIoU,未来帧预测达47.67 IoU。结果表明,高效稠密3D重建可超越传统感知目标,成为下一代自动驾驶的可扩展基础几何表征。
原文摘要 · Abstract (English)
Dense 3D reconstruction has demonstrated immense potential for spatial understanding, yet its viability as a real-time, onboard representation for autonomous driving remains an open challenge. Existing large-scale visual geometry models typically require substantial computational resources and lack the long-range geometric fidelity, surround-view consistency, and real-time efficiency demanded by dynamic driving environments. To bridge this gap, we present \textbf{LiAuto-GeoX}, an efficient grounded driving transformer designed for deployable, ego-centric 3D scene understanding. Our approach begins by learning a high-capacity driving geometry model from large-scale surround-view data, utilizing sparse LiDAR priors to provide robust geometric grounding in distant, ambiguous, or structure-sparse regions. We then instantiate this capability into a highly compact 155M-parameter onboard model through a novel geometry-preserving distillation framework. This framework employs mask-guided depth-aware distillation to retain fine-grained metric structures by emphasizing geometrically informative regions, and relative-pose relational distillation to enforce cross-view spatial consistency through pose-induced geometric relations. Extensive evaluations reveal that \textbf{LiAuto-GeoX} runs at 220 FPS on KITTI while maintaining high-fidelity dense reconstruction, enabling real-time deployment. The learned geometry transfers seamlessly to downstream autonomy tasks, achieving 90.6 PDMS in trajectory prediction, 24.63 mIoU in occupancy prediction, and 47.67 IoU in future-frame prediction. These all demonstrate that efficient dense 3D reconstruction can transcend its traditional role as a perception target to serve as a scalable, foundational geometric representation for next-generation autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。