arXiv:2506.12251cs.CVcs.LG2025-06被引 6

用三平面技术高效融合多摄像头数据,提升自动驾驶端到端模型推理速度。

Efficient Multi-Camera Tokenization with Triplanes for End-to-End Driving

  • 基于三平面的多视角图像编码方法,与相机数量和分辨率无关。
  • 相比传统图像块法减少72%令牌数,推理速度提升50%。
  • 适合嵌入式车载系统,兼顾高精度与实时性,适用于自动驾驶仿真测试。

自回归Transformer因其可扩展性和利用互联网规模预训练以实现泛化的能力,正被越来越多地用于端到端机器人和自动驾驶车辆(AV)策略架构。因此,高效地对传感器数据进行标记是确保此类架构在嵌入式硬件上实时可行的关键。为此,我们提出一种基于三平面的多摄像头高效标记策略,利用近期在3D神经重建与渲染方面的进展,生成与输入摄像头数量和分辨率无关、同时显式考虑车辆周围几何结构的传感器标记。在大规模自动驾驶数据集和最先进的神经模拟器上的实验表明,该方法相比当前基于图像块的标记策略有显著节省,最多减少72%的标记数,使策略推理速度最快提升50%,且保持相同的开环运动规划精度,并在闭环驾驶模拟中实现更高的离路率表现。

原文摘要 · Abstract (English)

Autoregressive Transformers are increasingly being deployed as end-to-end robot and autonomous vehicle (AV) policy architectures, owing to their scalability and potential to leverage internet-scale pretraining for generalization. Accordingly, tokenizing sensor data efficiently is paramount to ensuring the real-time feasibility of such architectures on embedded hardware. To this end, we present an efficient triplane-based multi-camera tokenization strategy that leverages recent advances in 3D neural reconstruction and rendering to produce sensor tokens that are agnostic to the number of input cameras and their resolution, while explicitly accounting for their geometry around an AV. Experiments on a large-scale AV dataset and state-of-the-art neural simulator demonstrate that our approach yields significant savings over current image patch-based tokenization strategies, producing up to 72% fewer tokens, resulting in up to 50% faster policy inference while achieving the same open-loop motion planning accuracy and improved offroad rates in closed-loop driving simulations.

自动驾驶三平面端到端高效标记

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。