构建可扩展的多跳推理视觉问答数据集,挑战现有模型理解能力
ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering
- 基于百科结构化知识自动生成复杂多跳问题
- 当前版本规模超现有知识型数据集一个数量级
- 适合评估模型深层推理与知识融合能力
本文提出新数据集 ReasonVQA,用于视觉问答任务。该数据集通过低成本框架自动整合结构化百科知识,生成复杂的多跳问题。在 ReasonVQA 上评估主流 VQA 模型,结果表明其对模型构成显著挑战,凸显其作为基准的潜力。此外,该数据集可轻松扩展图像输入,当前版本规模超过现有需外部知识的数据集一个数量级。
原文摘要 · Abstract (English)
In this paper, we propose a new dataset, ReasonVQA, for the Visual Question Answering (VQA) task. Our dataset is automatically integrated with structured encyclopedic knowledge and constructed using a low-cost framework, which is capable of generating complex, multi-hop questions. We evaluated state-of-the-art VQA models on ReasonVQA, and the empirical results demonstrate that ReasonVQA poses significant challenges to these models, highlighting its potential for benchmarking and advancing the field of VQA. Additionally, our dataset can be easily scaled with respect to input images; the current version surpasses the largest existing datasets requiring external knowledge by more than an order of magnitude.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。