arXiv:2602.00993cs.ROcs.AI2026-02被引 5

让自动驾驶更懂罕见危险场景,提升复杂路况下的安全决策能力。

HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving

  • 用大模型生成结构化风险提示,指导驾驶轨迹规划
  • 在真实长尾数据集上显著优于现有端到端方案
  • 适合研究自动驾驶安全与多模态融合的从业者

端到端自动驾驶模型越来越多地借助大视觉-语言模型提升语义理解能力,但在长尾条件下确保安全准确运行仍具挑战。这一挑战在长尾混合交通场景中尤为突出,自动驾驶车辆需在复杂不确定环境下与不同类型的路权使用者(包括人类驾驶车辆和弱势道路使用者)交互。本文提出HERMES,一种全栈式风险感知端到端多模态驾驶框架,通过注入显式的长尾风险信号来增强轨迹规划。HERMES采用基础模型辅助的标注流程,生成结构化的长尾场景上下文与长尾规划上下文,捕捉以隐患为中心的线索、操作意图及安全偏好,并利用这些信号引导端到端规划。此外,引入三模态驾驶模块,融合多视角感知、历史运动信息与语义引导,实现长尾场景下的风险感知精准轨迹规划。在真实世界长尾数据集上的实验表明,HERMES在长尾混合交通场景中持续优于代表性端到端及基于VLM的基线方法。消融实验证实了关键组件的互补贡献。

原文摘要 · Abstract (English)

End-to-end autonomous driving models increasingly benefit from large vision--language models for semantic understanding, yet ensuring safe and accurate operation under long-tail conditions remains challenging. These challenges are particularly prominent in long-tail mixed-traffic scenarios, where autonomous vehicles must interact with heterogeneous road users, including human-driven vehicles and vulnerable road users, under complex and uncertain conditions. This paper proposes HERMES, a holistic risk-aware end-to-end multimodal driving framework designed to inject explicit long-tail risk cues into trajectory planning. HERMES employs a foundation-model-assisted annotation pipeline to produce structured Long-Tail Scene Context and Long-Tail Planning Context, capturing hazard-centric cues together with maneuver intent and safety preference, and uses these signals to guide end-to-end planning. HERMES further introduces a Tri-Modal Driving Module that fuses multi-view perception, historical motion cues, and semantic guidance, ensuring risk-aware accurate trajectory planning under long-tail scenarios. Experiments on the real-world long-tail dataset demonstrate that HERMES consistently outperforms representative end-to-end and VLM-driven baselines under long-tail mixed-traffic scenarios. Ablation studies verify the complementary contributions of key components.

自动驾驶多模态风险感知长尾场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。