arXiv:2506.09557cs.CVcs.AI2025-06IJCV被引 6

提出新基准与方法,让AI在复杂城市场景下推理更可信。

Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes

  • 用分层证据链诊断模型推理过程
  • 发现定位与理解错误源于空间模糊和语义不确定
  • 通过显式生成证据单元提升可靠性,适合视觉推理研究者

尽管多模态大模型在简单场景表现优异,但在复杂城市环境下其认知可靠性显著下降。模型常依赖隐式推理而缺乏充分视觉证据,导致感知与推理脱节。现有评估仅关注最终结果,无法诊断推理缺陷。为此,作者提出AD2-Bench,引入分层视觉诊断框架,将推理分解为结构化证据链(CoE)。细粒度分析表明,鲁棒多模态推理关键在于准确获取证据。从概率视角出发,识别出两类推理失败原因:空间模糊(目标与背景混淆导致定位错误)和语义不确定(视觉特征退化引发理解错误)。为克服证据不足,提出基于证据的视觉推理(EGVOR),以显式生成结构化的时空-语义三元组(证据原子),强制对齐定位与理解。模型通过从反射监督构建到强化学习的分层课程训练,明确奖励降低推理方差。大量实验显示,EGVOR在恶劣条件下显著提升推理稳定性,为可信多模态认知提供更稳健框架。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes under adverse conditions. In these settings, models often rely on implicit inference without sufficient visual evidence, leading to a disconnect between perception and reasoning. Meanwhile, existing outcome-oriented benchmarks evaluate only final predictions and fail to diagnose failures in the underlying reasoning process. To address this gap, the authors propose AD2-Bench, which introduces a Hierarchical Visual Diagnosis framework that decomposes reasoning into a structured Chain of Evidence (CoE). This fine-grained diagnosis reveals that robust multimodal reasoning fundamentally depends on accurate evidence acquisition. Building on this perspective, the authors formulate reasoning from a probabilistic viewpoint and identify two primary causes of reasoning failure: Spatial Ambiguity, where models fail to distinguish target objects from background clutter, resulting in localization errors; and Semantic Uncertainty, where degraded visual features lead to incorrect semantic interpretation, resulting in understanding errors. To overcome these evidence deficiencies, they further propose Evidence-grounded Visual Reasoning (EGVOR), which replaces implicit reasoning with the explicit generation of Evidence Atoms - structured spatial-semantic triplets that enforce tight alignment between localization and semantic understanding. The model is trained through a hierarchical curriculum that progresses from reflective supervision construction to reinforcement learning, where reducing reasoning variance is explicitly rewarded. Extensive experiments demonstrate that EGVOR substantially improves reasoning stability under adverse conditions, providing a more robust framework for trustworthy multimodal cognition.

多模态推理可信AI视觉诊断城市场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。