arXiv:2503.19120cs.CLcs.AI2025-03NAACL被引 3

提出新评估方法,让文档VQA模型回答更贴近原文内容。

Where is this coming from? Making groundedness count in the evaluation of Document VQA models

  • 根据答案与文档的语义和多模态位置关系评分
  • 新方法能更好区分幻觉与正确答案,得分更合理
  • 适合关注模型可信度的研究者和开发者

近年来,文档视觉问答(Document VQA)模型进展迅速,部分基准上已接近甚至达到人类水平。然而,现有主流评估指标未考虑模型输出的语义和多模态根基性,导致幻觉与准确回答得分相同,无法反映模型的真实推理能力。为此,我们提出一种新型评估方法,结合输出的语义特征及其在输入文档中的多模态位置,量化其根基性。该方法参数化设计,可按用户偏好调整权重。通过人工判断验证,新方法能有效影响现有排行榜。大量分析表明,该方法更准确反映模型鲁棒性,更奖励校准良好的答案。

原文摘要 · Abstract (English)

Document Visual Question Answering (VQA) models have evolved at an impressive rate over the past few years, coming close to or matching human performance on some benchmarks. We argue that common evaluation metrics used by popular benchmarks do not account for the semantic and multimodal groundedness of a model's outputs. As a result, hallucinations and major semantic errors are treated the same way as well-grounded outputs, and the evaluation scores do not reflect the reasoning capabilities of the model. In response, we propose a new evaluation methodology that accounts for the groundedness of predictions with regard to the semantic characteristics of the output as well as the multimodal placement of the output within the input document. Our proposed methodology is parameterized in such a way that users can configure the score according to their preferences. We validate our scoring methodology using human judgment and show its potential impact on existing popular leaderboards. Through extensive analyses, we demonstrate that our proposed method produces scores that are a better indicator of a model's robustness and tends to give higher rewards to better-calibrated answers.

文档VQA评估方法模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。