为科学图表设计新VQA数据集,强调图表与原始数据的非一一对应关系。
What Lies Beneath: A Call for Distribution-based Visual Question & Answer Datasets
- 基于真实数据生成合成柱状图,模拟图表对数据的转化过程。
- 人类与大模型在无原始数据时答错率超60%,凸显数据依赖性。
- 适合研究多模态推理、科学图表理解的学者使用。
视觉问答(VQA)已成为评估大型多模态模型(LMMs)图像理解能力的重要基准。然而,现有VQA数据集多聚焦于真实世界图像或简单图表分析,极少关注复杂科学图表。许多图表类数据集不提供图表背后的原始数据,或假设图表标记与数据间存在一对一映射。事实上,图表是数据经分析、简化、修改后的产物。这一差异带来了当前数据集未捕捉的推理挑战。本文主张建立专门针对科学图表的VQA基准,其中图表标记与原始数据不存在一一对应关系。我们调研了现有数据集并指出其局限性,随后基于真实数据生成合成柱状图,并向人类和大型推理模型提出需依赖原始数据才能精确回答的问题。我们开源该数据集,包含图表、原始数据、生成数据的分布参数及所有图表标记与文本的边界框,供未来研究使用。
原文摘要 · Abstract (English)
Visual Question Answering (VQA) has become an important benchmark for assessing how large multimodal models (LMMs) interpret images. However, most VQA datasets focus on real-world images or simple diagrammatic analysis, with few focused on interpreting complex scientific charts. Indeed, many VQA datasets that analyze charts do not contain the underlying data behind those charts or assume a 1-to-1 correspondence between chart marks and underlying data. In reality, charts are transformations (i.e. analysis, simplification, modification) of data. This distinction introduces a reasoning challenge in VQA that the current datasets do not capture. In this paper, we argue for a dedicated VQA benchmark for scientific charts where there is no 1-to-1 correspondence between chart marks and underlying data. To do so, we survey existing VQA datasets and highlight limitations of the current field. We then generate synthetic histogram charts based on ground truth data, and ask both humans and a large reasoning model questions where precise answers depend on access to the underlying data. We release the open-source dataset, including figures, underlying data, distribution parameters used to generate the data, and bounding boxes for all figure marks and text for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。