构建多模态问答新基准,评估模型在复杂场景下的鲁棒性表现。
DARE: Diverse Visual Question Answering with Robustness Evaluation
- 设计五类多样化的视觉问答任务,涵盖计数与空间推理等难点。
- 发现顶尖模型在不同评测变体下性能下降最高达34%。
- 开源模型鲁棒性弱于闭源模型,但两者均对指令微调敏感。
视觉语言模型(VLMs)融合了文本与视觉模态的能力,能处理多模态输入。尽管在标准图像分类与图文匹配任务上表现良好,但在计数、空间推理等关键视觉语言推理能力上仍显不足。此外,现有基准未能充分评估模型对指令或评测协议微小变化的鲁棒性。为此,我们提出DARE(Diverse Visual Question Answering with Robustness Evaluation),一个精心设计的多项选择型视觉问答基准。DARE覆盖五类多样化场景,并包含四类鲁棒性评估:提示变化、答案选项子集变化、输出格式变化及正确答案数量变化。实验表明,当前最先进的VLM在多数类别中仍表现不佳,且在不同鲁棒性测试中无法稳定发挥最佳性能。答案选项子集变化导致最差情况性能较标准情况下降高达34%。开源模型如LLaVA 1.6和Idefics2的鲁棒性不及GPT-4和Gemini等闭源模型,但后者同样对各类变化极为敏感。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) extend remarkable capabilities of text-only large language models and vision-only models, and are able to learn from and process multi-modal vision-text input. While modern VLMs perform well on a number of standard image classification and image-text matching tasks, they still struggle with a number of crucial vision-language (VL) reasoning abilities such as counting and spatial reasoning. Moreover, while they might be very brittle to small variations in instructions and/or evaluation protocols, existing benchmarks fail to evaluate their robustness (or rather the lack of it). In order to couple challenging VL scenarios with comprehensive robustness evaluation, we introduce DARE, Diverse Visual Question Answering with Robustness Evaluation, a carefully created and curated multiple-choice VQA benchmark. DARE evaluates VLM performance on five diverse categories and includes four robustness-oriented evaluations based on the variations of: prompts, the subsets of answer options, the output format and the number of correct answers. Among a spectrum of other findings, we report that state-of-the-art VLMs still struggle with questions in most categories and are unable to consistently deliver their peak performance across the tested robustness evaluations. The worst case performance across the subsets of options is up to 34% below the performance in the standard case. The robustness of the open-source VLMs such as LLaVA 1.6 and Idefics2 cannot match the closed-source models such as GPT-4 and Gemini, but even the latter remain very brittle to different variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。