用视觉语言模型融合人车意图,实现更智能的自动驾驶共治。
SAFe-Copilot: Unified Shared Autonomy Framework
- 通过视觉语言模型理解驾驶意图与环境上下文,实现高层语义层面共治
- 在模拟测试中召回率100%,人类评估92%认同决策结果
- 在Bench2Drive上碰撞率显著下降,适合需人机协同的自动驾驶场景
自动驾驶系统在罕见、模糊和分布外场景下仍表现脆弱,而人类驾驶员能通过情境推理应对。共享自治通过在自主性不确定时引入人类输入来缓解此类问题。然而,现有方法多局限于低层轨迹仲裁,仅表达几何路径,无法保留驾驶意图。本文提出统一的共享自治框架,在更高抽象层次整合人类输入与自主规划。该方法利用视觉语言模型(VLM)从多模态线索(如驾驶行为与环境上下文)中推断驾驶员意图,并生成协调人机控制的一致策略。在模拟人类设定下,系统实现100%召回率,同时保持高准确率与精确度;人类受试者调查表明,92%的仲裁结果获得认同。在Bench2Drive基准上的评估显示,相比纯自主系统,碰撞率大幅降低,整体性能显著提升。以语义化、基于语言的表征进行仲裁,成为共享自治的设计原则,使系统具备常识推理能力并维持与人类意图的连续性。
原文摘要 · Abstract (English)
Autonomous driving systems remain brittle in rare, ambiguous, and out-of-distribution scenarios, where human driver succeed through contextual reasoning. Shared autonomy has emerged as a promising approach to mitigate such failures by incorporating human input when autonomy is uncertain. However, most existing methods restrict arbitration to low-level trajectories, which represent only geometric paths and therefore fail to preserve the underlying driving intent. We propose a unified shared autonomy framework that integrates human input and autonomous planners at a higher level of abstraction. Our method leverages Vision Language Models (VLMs) to infer driver intent from multi-modal cues -- such as driver actions and environmental context -- and to synthesize coherent strategies that mediate between human and autonomous control. We first study the framework in a mock-human setting, where it achieves perfect recall alongside high accuracy and precision. A human-subject survey further shows strong alignment, with participants agreeing with arbitration outcomes in 92% of cases. Finally, evaluation on the Bench2Drive benchmark demonstrates a substantial reduction in collision rate and improvement in overall performance compared to pure autonomy. Arbitration at the level of semantic, language-based representations emerges as a design principle for shared autonomy, enabling systems to exercise common-sense reasoning and maintain continuity with human intent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。