测试视觉模型在无法回答的问题上是否能拒绝作答
VisionTrap: Unanswerable Questions On Visual Data
- 构建VisionTrap数据集,含三类不可回答问题
- 模型在虚构图像前错误回答率超70%
- 适合评估模型的自我认知与可靠性
视觉问答(VQA)是广泛研究的领域,现有工作多关注模型对真实图像中可回答问题的响应。然而,对模型在面对不可回答问题时的表现,尤其是是否应主动回避作答,研究仍有限。本文提出新数据集VisionTrap,包含三类跨图像类型的不可回答问题:(1)融合物体与动物的混合实体;(2)以非现实或不可能场景呈现的物体;(3)虚构或不存在的人物。这些问题逻辑严谨但本质无解,旨在检验模型能否识别自身知识边界。实验结果表明,多数视觉语言模型在这些情境下仍强行生成答案,错误回答率超过70%。研究强调,将此类问题纳入基准评测至关重要,有助于评估模型是否具备真正的“不知”能力。
原文摘要 · Abstract (English)
Visual Question Answering (VQA) has been a widely studied topic, with extensive research focusing on how VLMs respond to answerable questions based on real-world images. However, there has been limited exploration of how these models handle unanswerable questions, particularly in cases where they should abstain from providing a response. This research investigates VQA performance on unrealistically generated images or asking unanswerable questions, assessing whether models recognize the limitations of their knowledge or attempt to generate incorrect answers. We introduced a dataset, VisionTrap, comprising three categories of unanswerable questions across diverse image types: (1) hybrid entities that fuse objects and animals, (2) objects depicted in unconventional or impossible scenarios, and (3) fictional or non-existent figures. The questions posed are logically structured yet inherently unanswerable, testing whether models can correctly recognize their limitations. Our findings highlight the importance of incorporating such questions into VQA benchmarks to evaluate whether models tend to answer, even when they should abstain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。