arXiv:2605.10564cs.CVcs.RO2026-05

通过预测未来帧潜在语义特征,实现端到端自动驾驶的长时序世界建模。

DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving

论文配图:DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving
图 1 · 摘自论文原文
  • 并行预测鸟瞰图空间中连续未来帧的潜在特征,构建长时序世界模型。
  • 在Bench2drive闭环基准上达到当前最优性能,尤其提升长尾场景表现。
  • 引入自适应文本推理机制,融合社会常识增强复杂场景决策能力。

端到端自动驾驶系统正越来越多地整合视觉语言模型(VLM)架构,利用文本或视觉推理提升驾驶决策的鲁棒性与准确性。然而,现有方法中的推理机制大多直接来自通用领域,缺乏针对自动驾驶场景的深入设计,尤其在视觉推理模块方面存在不足。本文提出一种驾驶世界模型,可在鸟瞰图(BEV)空间中并行预测连续未来帧的潜在语义特征,从而实现对长期未来世界状态的建模。同时,我们引入一种高效且自适应的文本推理机制,利用额外的社会知识与推理能力,在挑战性长尾场景中进一步提升驾驶性能。所提方法在闭环基准Bench2drive上取得当前最优(SOTA)结果。代码已开源:https://github.com/hotdogcheesewhite/DeepSight。

原文摘要 · Abstract (English)

End-to-end autonomous driving systems are increasingly integrating Vision-Language Model (VLM) architectures, incorporating text reasoning or visual reasoning to enhance the robustness and accuracy of driving decisions. However, the reasoning mechanisms employed in most methods are direct adaptations from general domains, lacking in-depth exploration tailored to autonomous driving scenarios, particularly within visual reasoning modules. In this paper, we propose a driving world model that performs parallel prediction of latent semantic features for consecutive future frames in the bird's-eye-view (BEV) space, thereby enabling long-horizon modeling of future world states. We also introduce an efficient and adaptive text reasoning mechanism that utilizes additional social knowledge and reasoning capabilities to further improve driving performance in challenging long-tail scenarios. We present a novel, efficient, and effective approach that achieves state-of-the-art (SOTA) results on the closed-loop Bench2drive benchmark. Codes are available at: https://github.com/hotdogcheesewhite/DeepSight.

自动驾驶世界建模视觉语言模型长时序预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。