arXiv:2506.21041cs.ROcs.AI2025-06被引 1

用视觉语言模型提升极端天气下自动驾驶的安全性

SEAL: Vision-Language Model-Based Safe End-to-End Cooperative Autonomous Driving with Adaptive Long-Tail Modeling

  • 基于提示生成罕见天气场景,增强训练多样性
  • 自适应注意力模块根据场景修正模糊视觉特征
  • 多任务对比学习提升跨场景特征区分能力

自动驾驶在罕见、多样且视觉退化的天气条件下面临重大安全挑战,尤其在车路协同场景中更为突出。为此,我们提出SEAL框架,基于视觉-语言模型实现鲁棒的协同自动驾驶。其核心创新包括:(i) 利用基础模型驱动的提示式长尾场景生成与评估管道,合成雪、雾等真实长尾条件下的车端与路端视图,高效丰富训练数据;(ii) 门控多场景自适应注意力模块,利用场景先验对视觉流进行调制,重校准模糊或受损特征;(iii) 多任务场景感知对比学习目标,增强模态对齐并促进跨场景特征可分性。大量实验表明,SEAL在复杂驾驶条件下显著优于现有基线,在推理、安全与规划准确性上均取得提升,推动了自动驾驶的安全性、鲁棒性与可扩展性。

原文摘要 · Abstract (English)

Autonomous driving technologies face significant safety challenges while operating under rare, diverse, and visually degraded weather scenarios. These challenges become more critical in cooperative settings, where vehicles and infrastructure jointly perceive and reason across complex environments. To address these issues, we propose SEAL, a vision-language model-based framework with adaptive multimodal learning for robust cooperative autonomous driving under long-tail scenarios. SEAL introduces three core innovations: (i) a prompt-driven long-tail scenario generation and evaluation pipeline that leverages foundation models to synthesize realistic long-tail conditions such as snow and fog across vehicle- and infrastructure-side views, enriching training diversity efficiently; (ii) a gated multi-scenario adaptive attention module that modulates the visual stream using scenario priors to recalibrate ambiguous or corrupted features; and (iii) a multi-task scenario-aware contrastive learning objective that improves multimodal alignment and promotes cross-scenario feature separability. Extensive experiments demonstrate that SEAL significantly outperforms existing baselines in reasoning, safety, and planning accuracy under complex, challenging driving conditions, advancing the safety, robustness, and scalability of autonomous driving.

自动驾驶视觉语言模型长尾分布车路协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。