arXiv:2602.11509cs.CLcs.AI2026-02中稿 · ICML被引 3

让多模态模型生成带精准引用的答案,提升事实可验证性。

Multimodal Fact-Level Attribution for Verifiable Reasoning

  • 设计新基准MuRGAt,要求模型对视频、音频等多模态输入进行推理并标注具体来源片段。
  • 发现强模型虽推理正确却常虚构引用,且复杂推理会降低准确率。
  • 适合关注可信AI、可解释生成与多模态验证的研究者。

多模态大语言模型(MLLMs)越来越多用于涉及多步推理和长文本生成的现实任务,其可靠性依赖于将输出基于异构输入源并验证具体事实。然而,现有多模态接地评估基准和方法仅聚焦简化观测场景或有限模态,无法评估复杂多模态推理中的事实溯源。我们提出MuRGAt(Multimodal Reasoning with Grounded Attribution),一个评估复杂多模态推理中事实级溯源能力的基准。该基准要求模型在包含视频、音频等多种模态的输入下,生成带有明确推理过程和精确引用的答案,每条引用需标明模态类型及时间片段。为实现可靠评估,我们引入与人工判断高度相关的自动评估框架。人类与自动评分的对比显示,即使强大的MLLM也频繁虚构引用,尽管推理正确。此外,我们观察到关键权衡:增加推理深度或强制结构化接地常导致准确率下降,揭示了内部推理与可验证溯源之间存在显著差距。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are increasingly used for real-world tasks involving multi-step reasoning and long-form generation, where reliability requires grounding model outputs in heterogeneous input sources and verifying individual factual claims. However, existing multimodal grounding benchmarks and evaluation methods focus on simplified, observation-based scenarios or limited modalities and fail to assess attribution in complex multimodal reasoning. We introduce MuRGAt (Multimodal Reasoning with Grounded Attribution), a benchmark for evaluating fact-level multimodal attribution in settings that require reasoning beyond direct observation. Given inputs spanning video, audio, and other modalities, MuRGAt requires models to generate answers with explicit reasoning and precise citations, where each citation specifies both modality and temporal segments. To enable reliable assessment, we introduce an automatic evaluation framework that strongly correlates with human judgments. Benchmarking with human and automated scores reveals that even strong MLLMs frequently hallucinate citations despite correct reasoning. Moreover, we observe a key trade-off: increasing reasoning depth or enforcing structured grounding often degrades accuracy, highlighting a significant gap between internal reasoning and verifiable attribution.

多模态事实验证可解释性推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。