评测模型能否在图表中定位支撑答案的视觉证据。
DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams

- 要求模型精准标注支持答案的图表区域,而非依赖文本关联
- 涵盖6个数据集共11,664个问题,测试集2,445个带人工验证标注
- 适合研究可解释性、视觉推理与模型可信度的学者使用
图表问答(DQA)要求模型理解图表、地图、信息图、电路图和科学图示等结构化视觉内容。尽管近期视觉语言模型(VLMs)在这些任务上表现优异,但高准确率并不意味着模型真正基于图表中的视觉证据进行推理。模型可能仅依赖文本关联或数据集固有偏差,未正确识别支撑答案的视觉元素。这限制了对图表推理能力的真实评估并降低可解释性。为此,我们提出DRAGON,一个用于评估图表中证据接地推理的基准。给定图表、问题和正确答案,模型需预测出支撑答案的视觉区域边界框,包括答案组件、文本标签、图例、坐标轴、连接线及其他推理相关结构。DRAGON数据集包含来自六个图表问答数据集(ChartQA、Circuit-VQA、InfographicsVQA、MapIQ、MapWise、AI2D)的11,664个标注问题实例。我们发布了一个2,445个实例的测试集,附有人工验证的推理证据标注,并提供标准化评估框架。我们评估了八种最新VLMs在不同图表领域中定位推理证据的能力。DRAGON推动了图表推理的系统性评估,支持未来致力于视觉证据接地模型的研究。
原文摘要 · Abstract (English)
Diagram question answering (DQA) requires models to interpret structured visual representations such as charts, maps, infographics, circuit schematics, and scientific diagrams. Recent vision-language models (VLMs) often achieve high answer accuracy on these tasks, yet correct answers do not guarantee that models ground their reasoning in the diagram regions that support the prediction. Models may instead rely on textual correlations or dataset artifacts without identifying the visual evidence required to verify the answer. This limitation prevents reliable evaluation of diagram reasoning and reduces interpretability. We introduce DRAGON, a benchmark for evaluating evidence-grounded visual reasoning in diagrams. Given a diagram, a question, and the correct answer, a model must predict bounding boxes that correspond to the visual elements required to justify the answer. These evidence regions may include answer-bearing components, textual labels, legends, axes, connectors, and other supporting structures involved in the reasoning process. The DRAGON dataset contains 11,664 annotated question instances collected from six diagram QA datasets: ChartQA, Circuit-VQA, InfographicsVQA, MapIQ, MapWise, and AI2D. We release a 2,445-instance benchmark test set with human-verified reasoning evidence annotations and a standardized evaluation framework. We evaluate eight recent VLMs and analyze their ability to localize reasoning evidence across diverse diagram domains. DRAGON enables systematic evaluation of diagram reasoning and supports future research on models that ground their predictions in visual evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。