新基准测试揭示大模型在图文问答中事实错误频发,且视觉与语言模块均存短板。
VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering
- 分离评估图文模型的视觉与语言能力,精准定位弱点。
- 顶尖模型GPT-4o在困难集上仅30%正确率,暴露严重缺陷。
- 适合研究多模态模型可靠性与可解释性的学者使用。
大型视觉语言模型(LVLM)在图文问答任务中虽表现优异,但生成非事实性回答的现象仍普遍存在。现有多模态事实类问答评测主要依赖模型输出与真实答案的对比,难以深入分析各模态模块的表现。为此,我们提出VisualSimpleQA,一个支持解耦评估的多模态事实问答基准,具备两大特点:一是实现对视觉与语言模态的独立、高效评估;二是引入明确的难度标准,指导人工标注并提取出更具挑战性的子集——VisualSimpleQA-hard。对15个LVLM的实验表明,即使最先进的模型如GPT-4o,在VisualSimpleQA上的准确率也仅达60%以上,在VisualSimpleQA-hard上更是低于30%。解耦评估结果揭示了视觉与语言模块均有巨大改进空间。数据集已公开于https://huggingface.co/datasets/WYLing/VisualSimpleQA。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) have demonstrated remarkable achievements, yet the generation of non-factual responses remains prevalent in fact-seeking question answering (QA). Current multimodal fact-seeking benchmarks primarily focus on comparing model outputs to ground truth answers, providing limited insights into the performance of modality-specific modules. To bridge this gap, we introduce VisualSimpleQA, a multimodal fact-seeking benchmark with two key features. First, it enables streamlined and decoupled evaluation of LVLMs in visual and linguistic modalities. Second, it incorporates well-defined difficulty criteria to guide human annotation and facilitates the extraction of a challenging subset, VisualSimpleQA-hard. Experiments on 15 LVLMs show that even state-of-the-art models such as GPT-4o achieve merely 60%+ correctness in multimodal fact-seeking QA on VisualSimpleQA and 30%+ on VisualSimpleQA-hard. Furthermore, the decoupled evaluation across these models highlights substantial opportunities for improvement in both visual and linguistic modules. The dataset is available at https://huggingface.co/datasets/WYLing/VisualSimpleQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。