arXiv:2507.12490cs.CVcs.AI2025-07中稿 · presentation at th…

无需微调即可生成可定位的文档问答解释,提升透明度与可复现性。

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering

  • 通过视觉语言模型生成自然语言推理,再用多模态嵌入相似度定位到图像区域。
  • 在DocVQA上精确匹配准确率和归一化编辑距离均优于基线模型。
  • 完全免训练、通用性强,适合需要可解释性的文档智能应用。

我们提出EaGERS,一个完全无需训练且模型无关的流程:(1) 通过视觉语言模型生成自然语言推理;(2) 利用可配置网格上的多模态嵌入相似度与多数投票机制,将推理结果空间定位到图像子区域;(3) 仅允许从掩码后选定的相关区域生成回答。在DocVQA数据集上的实验表明,最佳配置不仅在精确匹配准确率和平均归一化莱文斯坦相似度指标上超越基线模型,还提升了文档视觉问答任务中的透明度与可复现性,且无需额外模型微调。

原文摘要 · Abstract (English)

We introduce EaGERS, a fully training-free and model-agnostic pipeline that (1) generates natural language rationales via a vision language model, (2) grounds these rationales to spatial sub-regions by computing multimodal embedding similarities over a configurable grid with majority voting, and (3) restricts the generation of responses only from the relevant regions selected in the masked image. Experiments on the DocVQA dataset demonstrate that our best configuration not only outperforms the base model on exact match accuracy and Average Normalized Levenshtein Similarity metrics but also enhances transparency and reproducibility in DocVQA without additional model fine-tuning.

可解释AI文档理解视觉问答空间定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。