提出显式证据生成机制,提升复杂城市场景下多模态推理的可信度。
Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
- 通过结构化证据链分解推理过程,实现细粒度诊断
- 在恶劣条件下推理准确率显著提升,定位与理解错误减少
- 适合关注多模态系统可靠性与可解释性的研究者
尽管多模态大模型在理想场景中表现优异,但在复杂城市场景的不利条件下,其认知可靠性明显下降。模型常依赖隐式推断而缺乏充分视觉证据,导致感知与推理脱节。现有以结果为导向的评估基准仅关注最终预测,无法诊断推理过程中的失败。为此,作者提出AD2-Bench,引入分层视觉诊断框架,将推理分解为结构化的证据链(CoE)。该诊断揭示:稳健的多模态推理根本依赖于准确的证据获取。基于此,作者从概率视角出发,识别出两类推理失败原因:空间模糊性(模型难以区分目标与背景,导致定位错误)和语义不确定性(视觉特征退化导致语义误判)。为克服证据缺失问题,提出证据根基的视觉推理(EGVOR),将隐式推理替换为显式生成‘证据原子’——即结构化的空间-语义三元组,强制实现定位与理解的一致性。模型通过从反思性监督构建到强化学习的分层训练流程进行训练,明确奖励降低推理方差。大量实验表明,EGVOR在恶劣条件下显著提升推理稳定性,为可信多模态认知提供更鲁棒的框架。
原文摘要 · Abstract (English)
While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes under adverse conditions. In these settings, models often rely on implicit inference without sufficient visual evidence, leading to a disconnect between perception and reasoning. Meanwhile, existing outcome-oriented benchmarks evaluate only final predictions and fail to diagnose failures in the underlying reasoning process. To address this gap, the authors propose AD2-Bench, which introduces a Hierarchical Visual Diagnosis framework that decomposes reasoning into a structured Chain of Evidence (CoE). This fine-grained diagnosis reveals that robust multimodal reasoning fundamentally depends on accurate evidence acquisition. Building on this perspective, the authors formulate reasoning from a probabilistic viewpoint and identify two primary causes of reasoning failure: Spatial Ambiguity, where models fail to distinguish target objects from background clutter, resulting in localization errors; and Semantic Uncertainty, where degraded visual features lead to incorrect semantic interpretation, resulting in understanding errors. To overcome these evidence deficiencies, they further propose Evidence-grounded Visual Reasoning (EGVOR), which replaces implicit reasoning with the explicit generation of Evidence Atoms - structured spatial-semantic triplets that enforce tight alignment between localization and semantic understanding. The model is trained through a hierarchical curriculum that progresses from reflective supervision construction to reinforcement learning, where reducing reasoning variance is explicitly rewarded. Extensive experiments demonstrate that EGVOR substantially improves reasoning stability under adverse conditions, providing a more robust framework for trustworthy multimodal cognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。