arXiv:2607.01420cs.CLcs.AI2026-07

无需训练即可精准定位长文档多模态答案来源,提升可信度与效率。

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering

论文配图:MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering
图 1 · 摘自论文原文
  • 利用预填充阶段注意力头和校准阈值实现无训练归因
  • 在多模态长文档上达到优于提示法和GPT 5.4的归因准确率
  • 推理延迟仅为提示法的1/7,适合高实时性场景

随着基于证据的问答系统在AI助手中的广泛应用,将生成的答案准确归因于证据对用户信任和模型安全至关重要。尽管单模态归因已深入研究,多模态设置仍相对缺乏。为此,我们提出MultAttnAttrib,一种无需训练的归因生成方法,利用模型预填充过程、选定注意力头和校准阈值,在文档中定位源证据。为建立方法基准,我们引入MultAttrEval,一个标注了细粒度真实归因的互补基准数据集,专为长文档多模态归因设计。据我们所知,这是首个针对长文档多模态归因的评估数据集。实验表明,MultAttnAttrib始终优于多种归因生成方法,包括若干强提示法,并可媲美最新前沿模型如GPT 5.4。该方法不仅显著提升单模态与多模态归因准确率,且在相同基模型上归因延迟低至提示法的七分之一。

原文摘要 · Abstract (English)

As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for user trust and model safety. While unimodal attributions have been explored in depth, the multimodal setting remains relatively under-researched. As a result, we introduce MultAttnAttrib, a training-free attribution-generation method that leverages a model's prefill pass, selected attention heads, and calibrated thresholds to locate source evidence within a document. To establish baseline results for the method, we introduce MultAttrEval, a complementary benchmark dataset annotated with fine-grained, ground-truth attributions for answer components grounded in multimodal source documents. To our knowledge, this is the first evaluation dataset designed specifically for multimodal attribution in long-form documents. Experimental results show that MultAttnAttrib consistently outperforms a variety of attribution-generation methods, including several strong prompting-based approaches and matches the latest frontier models such as GPT 5.4. Our method not only substantially improves attribution accuracy for both unimodal and multimodal attribution types, but also produces attributions at up to one-seventh of the direct inference latency compared to prompting on the same base model.

多模态归因长文档问答无训练方法可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。