让自动驾驶理解人类抽象意图,突破传统导航局限
From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving
- 构建意图驱动的端到端自动驾驶框架,融合视觉与语言信息
- 提出新评估方法IFA,从语义层面衡量意图实现程度
- 适合研究智能驾驶、人机交互及多模态决策的学者
尽管端到端自动驾驶在几何控制方面取得显著进展,现有系统仍受限于依赖简单指令的命令跟随范式。实现真正智能的自动驾驶代理,需具备解读并实现高层级抽象人类意图的能力。然而,这一进展因缺乏专用基准和语义感知评估指标而受阻。本文正式定义了意图驱动的端到端自动驾驶任务,并提出了Intention-Drive基准,构建了一个大规模数据集,包含复杂自然语言意图与高保真传感器数据。为克服传统轨迹评估指标的局限,我们引入想象未来对齐(IFA)评估协议,利用生成式世界模型评估目标语义实现程度,超越几何精度。此外,我们探索解决方案空间,提出两种不同范式:端到端视觉-语言规划器与分层代理框架。实验揭示现有模型虽有良好驾驶稳定性,但在意图实现上表现不佳;所提框架显著提升与人类意图的对齐度。
原文摘要 · Abstract (English)
While end-to-end autonomous driving has achieved remarkable progress in geometric control, current systems remain constrained by a command-following paradigm that relies on simple navigational instructions. Transitioning to genuinely intelligent agents requires the capability to interpret and fulfill high-level, abstract human intentions. However, this advancement is hindered by the lack of dedicated benchmarks and semantic-aware evaluation metrics. In this paper, we formally define the task of Intention-Driven End-to-End Autonomous Driving and present Intention-Drive, a comprehensive benchmark designed to bridge this gap. We construct a large-scale dataset featuring complex natural language intentions paired with high-fidelity sensor data. To overcome the limitations of conventional trajectory-based metrics, we introduce the Imagined Future Alignment (IFA), a novel evaluation protocol leveraging generative world models to assess the semantic fulfillment of human goals beyond mere geometric accuracy. Furthermore, we explore the solution space by proposing two distinct paradigms: an end-to-end vision-language planner and a hierarchical agent-based framework. The experiments reveal a critical dichotomy where existing models exhibit satisfactory driving stability but struggle significantly with intention fulfillment. Notably, the proposed frameworks demonstrate superior alignment with human intentions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。