arXiv:2607.23537cs.AI2026-07

构建真实恶劣天气下多模态自动驾驶评估基准,揭示模型可靠性下降根源。

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness

论文配图:ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness
图 1 · 摘自论文原文
  • 基于可观测性元标注构建多模态评测框架,涵盖感知、空间可靠性和风险决策三维度。
  • 在超14,000个训练与13,000个测试样本上验证,现有视觉语言模型性能显著下降。
  • 提出新模型通过正常与恶劣天气联合训练,提升复杂环境下的鲁棒性,适合自动驾驶研究者参考。

自动驾驶在恶劣天气下的表现仍是重大挑战,而现有视觉-语言基准多聚焦于标准条件、合成干扰或单模态评估。为此,我们指出关键难点在于环境可观测性下降:在雾、雨、雪及低光照条件下,多模态观测变得不可靠且跨模态不一致,影响场景理解与决策。为此,我们提出 extbf{ObsDriveBench}——一个面向真实恶劣天气的多模态自动驾驶基准。该基准设计包含三个能力维度: extbf{可观测性意识}、 extbf{空间可靠性} 和 extbf{风险感知决策},支持对模型行为的细粒度诊断。通过可观测性元标注、场景描述与面向能力的多选任务,融合同步摄像头、激光雷达和雷达输入,构建了包含超过14,000个训练样本和13,000个测试样本的数据集。实验表明,现有视觉-语言模型在恶劣天气下普遍存在性能退化。我们进一步提出 extbf{ObsDrive} 模型,结合正常天气监督微调与恶劣天气强化学习,显著提升了在三大能力上的鲁棒性。数据集与评估代码将开源至 exttt{ObsDriveBench}。

原文摘要 · Abstract (English)

Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality. As a result, it remains unclear how vision-language models behave under real-world adverse weather with multi-modal inputs. We argue that a key difficulty lies in degraded environmental observability: under fog, rain, snow, and low illumination, multi-modal observations become unreliable and cross-modally inconsistent, posing challenges to scene understanding, and subsequent decision-making. To study this, we introduce \textbf{ObsDriveBench}, a real-world multi-modal benchmark for adverse-weather autonomous driving. Our benchmark is designed with three capability dimensions: \textbf{observability awareness}, \textbf{spatial reliability}, and \textbf{risk-aware decision-making}, enabling fine-grained diagnosis of model behavior under degraded observations. We construct the benchmark through observability meta-annotation, scene description, and capability oriented multiple-choice tasks over synchronized camera, LiDAR, and radar inputs, forming a benchmark with over 14k training and 13k test questions. Experiments reveal consistent performance degradation of existing vision-language models. We further introduce \textbf{ObsDrive} model with normal-weather supervised fine-tuning and adverse-weather reinforcement learning, improving robustness across all three capabilities. The dataset and evaluation code will be released at \href{https://github.com/russellyq/ObsDriveBench}{\texttt{ObsDriveBench}}.

自动驾驶多模态恶劣天气评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。