构建《辛普森一家》动漫图像问答数据集,推动探究式学习AI研究
SimpsonsVQA: Enhancing Inquiry-Based Learning with a Tailored Dataset
- 基于动漫图像设计多任务问答框架,支持问-答、无关问题识别与答案评判
- 包含23万张图像、166万组问答对及50万条判断数据,规模超现有同类数据集
- 揭示大模型在卡通图像上零样本表现不足,适合教育类AI与交互系统研发
视觉问答(VQA)已成为推动交互式沉浸式学习的有前景方向。尽管已有众多VQA数据集用于回答问题或识别不可回答问题,但大多数基于真实世界图像,缺乏对动漫图像上模型性能的探索。为此,本文提出SimpsonsVQA,一个源自《辛普森一家》电视节目的新数据集,旨在促进探究式学习。该数据集不仅涵盖传统VQA任务,还包含识别与图像无关的问题,以及用户给出答案后系统评估其正确性、错误性或模糊性的逆向任务。数据集包含约23,000张图像、166,000组问答对和500,000条判断结果(https://simpsonsvqa.org)。实验表明,当前主流视觉语言模型如ChatGPT4o在零样本设置下三项任务均表现不佳,凸显该数据集在提升动漫图像理解能力方面的价值。我们期待SimpsonsVQA能激发更多关于探究式学习VQA的研究与创新。
原文摘要 · Abstract (English)
Visual Question Answering (VQA) has emerged as a promising area of research to develop AI-based systems for enabling interactive and immersive learning. Numerous VQA datasets have been introduced to facilitate various tasks, such as answering questions or identifying unanswerable ones. However, most of these datasets are constructed using real-world images, leaving the performance of existing models on cartoon images largely unexplored. Hence, in this paper, we present "SimpsonsVQA", a novel dataset for VQA derived from The Simpsons TV show, designed to promote inquiry-based learning. Our dataset is specifically designed to address not only the traditional VQA task but also to identify irrelevant questions related to images, as well as the reverse scenario where a user provides an answer to a question that the system must evaluate (e.g., as correct, incorrect, or ambiguous). It aims to cater to various visual applications, harnessing the visual content of "The Simpsons" to create engaging and informative interactive systems. SimpsonsVQA contains approximately 23K images, 166K QA pairs, and 500K judgments (https://simpsonsvqa.org). Our experiments show that current large vision-language models like ChatGPT4o underperform in zero-shot settings across all three tasks, highlighting the dataset's value for improving model performance on cartoon images. We anticipate that SimpsonsVQA will inspire further research, innovation, and advancements in inquiry-based learning VQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。