arXiv:2509.10129cs.CLcs.IR2025-09被引 3

让视觉语言模型更准定位文档答案,提升可解释性。

Towards Reliable and Interpretable Document Question Answering via VLMs

  • 用独立框选模块分离答案生成与位置定位。
  • 发现正确答案常无可靠位置标注,揭示定位短板。
  • 适配现有模型,尤其适合无法微调的商用系统。

视觉语言模型(VLMs)在文档理解方面表现出色,尤其擅长从复杂文档中识别和提取文本信息。然而,准确地定位答案在文档中的位置仍是重大挑战,限制了其可解释性和实际应用。为此,我们提出DocExplainerV0,一个即插即用的边界框预测模块,将答案生成与空间定位解耦。该设计可直接应用于现有VLMs,包括无法微调的专有系统。通过系统评估,我们量化分析了文本准确性与空间定位之间的差距,发现正确答案往往缺乏可靠的定位。我们的标准化框架揭示了这些缺陷,并为未来更具可解释性和鲁棒性的文档信息抽取VLMs建立了基准。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have shown strong capabilities in document understanding, particularly in identifying and extracting textual information from complex documents. Despite this, accurately localizing answers within documents remains a major challenge, limiting both interpretability and real-world applicability. To address this, we introduce DocExplainerV0, a plug-and-play bounding-box prediction module that decouples answer generation from spatial localization. This design makes it applicable to existing VLMs, including proprietary systems where fine-tuning is not feasible. Through systematic evaluation, we provide quantitative insights into the gap between textual accuracy and spatial grounding, showing that correct answers often lack reliable localization. Our standardized framework highlights these shortcomings and establishes a benchmark for future research toward more interpretable and robust document information extraction VLMs.

文档问答视觉语言模型可解释性定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。