让自动驾驶理解全局导航信息,突破视觉范围限制。
NavigScene: Bridging Local Perception and Global Navigation for Beyond-Visual-Range Autonomous Driving
- 用导航引导的语言数据增强视觉模型推理能力
- 在多个任务上显著提升感知与规划性能
- 适合研究智能驾驶多模态融合的开发者
自动驾驶系统在局部视觉信息下的感知、预测和规划已取得显著进展,但难以整合人类司机常用的全局导航上下文。为此,我们提出NavigScene——一个辅助导航引导的自然语言数据集,模拟人类驾驶环境。同时开发三种互补范式:(1) 导航引导推理,通过提示注入导航上下文增强视觉语言模型;(2) 导航引导偏好优化,基于强化学习扩展直接偏好优化,建立对导航相关摘要信息的偏好;(3) 导航引导视觉-语言-动作模型,通过特征融合将导航引导与视觉语言模型结合传统驾驶模型。大量实验表明,所提方法显著提升感知、预测、规划及问答任务表现,使系统具备超越视觉范围的推理能力和在多样场景中的泛化性。本工作推动了更全面、可靠且安全的自动驾驶系统发展。
原文摘要 · Abstract (English)
Autonomous driving systems have made significant advances in Q&A, perception, prediction, and planning based on local visual information, yet they struggle to incorporate broader navigational context that human drivers routinely utilize. We address this critical gap between local sensor data and global navigation information by proposing NavigScene, an auxiliary navigation-guided natural language dataset that simulates a human-like driving environment within autonomous driving systems. Moreover, we develop three complementary paradigms to leverage NavigScene: (1) Navigation-guided Reasoning, which enhances vision-language models by incorporating navigation context into the prompting approach; (2) Navigation-guided Preference Optimization, a reinforcement learning method that extends Direct Preference Optimization to improve vision-language model responses by establishing preferences for navigation-relevant summarized information; and (3) Navigation-guided Vision-Language-Action model, which integrates navigation guidance and vision-language models with conventional driving models through feature fusion. Extensive experiments demonstrate that our approaches significantly improve performance across perception, prediction, planning, and question-answering tasks by enabling reasoning capabilities beyond visual range and improving generalization to diverse driving scenarios. This work represents a significant step toward more comprehensive autonomous driving systems capable of navigating complex, unfamiliar environments with greater reliability and safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。