首个关注视觉问题焦点模糊性的数据集,提升VQA模型对多义区域的识别能力。
Acknowledging Focus Ambiguity in Visual Questions
- 构建首个标注问题可能指向图像多个区域的数据集
- 现代模型在识别模糊性和定位多区域上表现不佳
- 适合研究视觉理解中语义歧义的学者使用
现有的视觉问答(VQA)研究均未考虑问题所指内容在图像中的位置模糊性。为填补这一空白,我们提出VQ-FocusAmbiguity,这是首个将每个合理答案对应图像中所有可能区域进行视觉标注的VQA数据集。我们分析并对比该数据集与现有数据集的特性,揭示其独特性。进一步地,我们针对两个新任务对现代模型进行基准测试:识别问题是否存在焦点模糊性,以及定位图像中所有可能的焦点区域。实验结果表明,该数据集对当前模型仍具挑战性。为推动未来研究,我们已公开数据集及评测服务器(https://vizwiz.org/tasks-and-datasets/focus-ambiguity-in-visual-questions)。
原文摘要 · Abstract (English)
No published work on visual question answering (VQA) accounts for ambiguity regarding where the content described in the question is located in the image. To fill this gap, we introduce VQ-FocusAmbiguity, the first VQA dataset that visually grounds each plausible image region a question could refer to when arriving at valid answers. We next analyze and compare our dataset to existing datasets to reveal its unique properties. Finally, we benchmark modern models for two novel tasks related to acknowledging focus ambiguity: recognizing whether a visual question has focus ambiguity and locating all plausible focus regions within the image. Results show that the dataset is challenging for modern models. To facilitate future progress on these tasks, we publicly share the dataset with an evaluation server at https://vizwiz.org/tasks-and-datasets/focus-ambiguity-in-visual-questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。