arXiv:2512.10947cs.CV2025-12被引 1

用少量可学习的场景令牌联合编码多视角视频,提升自动驾驶推理效率与性能。

Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving

  • 用可学习的场景令牌统一编码多摄像头多帧图像,无需3D先验知识。
  • 在2万小时数据上实现2.2倍推理吞吐量提升,驾驶表现显著超越现有方法。
  • 模型自发产生场景分解能力,适合追求高效端到端自动驾驶系统的研究者。

我们提出Flex,一种高效且有效的场景编码器,解决端到端自动驾驶中处理海量多摄像头数据的计算瓶颈。Flex使用一组少量可学习的场景令牌,联合编码来自不同摄像头和时间步的所有图像令牌。设计上,该方法不依赖鸟瞰图(BEV)、占据网格或三平面等显式三维归纳偏置,而是直接从数据中学习紧凑的场景表示。这种整体编码策略大幅压缩视觉输入,便于下游基于大语言模型(LLM)的策略模型处理。在包含20,000小时驾驶数据的大型私有数据集上评估,Flex实现了2.2倍的推理吞吐量提升,同时驾驶性能显著优于当前最优方法。此外,这些紧凑的场景令牌展现出无需显式监督的场景分解能力。研究结果挑战了三维先验必要的主流假设,表明数据驱动的联合编码策略为未来自动驾驶系统提供了更可扩展、高效且有效的路径。

原文摘要 · Abstract (English)

We present Flex, an efficient and effective scene encoder that addresses the computational bottleneck of processing high-volume multi-camera data in end-to-end autonomous driving. Flex employs a small set of learnable scene tokens to jointly encode information from all image tokens across different cameras and timesteps. By design, our approach is geometry-agnostic, learning a compact scene representation directly from data without relying on the explicit 3D inductive biases, such as Bird-Eye-View (BEV), occupancy or tri-plane representations, which are common in prior work. This holistic encoding strategy aggressively compresses the visual input for the downstream Large Language Model (LLM) based policy model. Evaluated on a large-scale proprietary dataset of 20,000 driving hours, our Flex achieves 2.2x greater inference throughput while improving driving performance by a large margin compared to state-of-the-art methods. Furthermore, we show that these compact scene tokens develop an emergent capability for scene decomposition without any explicit supervision. Our findings challenge the prevailing assumption that 3D priors are necessary, demonstrating that a data-driven, joint encoding strategy offers a more scalable, efficient and effective path for future autonomous driving systems.

端到端自动驾驶多相机编码场景表示大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。