arXiv:2502.09818cs.CV2025-02被引 11

测试视觉语言模型在科学问答中对图文干扰的鲁棒性

On the robustness of multimodal language model towards distractions

  • 构建包含图文干扰的新基准评估模型抗干扰能力
  • 多数顶级模型(如GPT-4)在干扰下推理性能显著下降
  • 文本干扰比图像干扰影响更大,提示工程可部分缓解

尽管视觉语言模型(VLMs)在视觉问答等任务中取得显著进展,其对提示变化的鲁棒性仍缺乏深入研究。理解干扰对VLMs的影响对提升其实际应用至关重要,因真实场景中输入常含噪声和无关信息。本文旨在评估VLMs在科学问答任务中面对图文干扰时的鲁棒性。基于ScienceQA数据集,我们构建了新基准,引入视觉与文本双重干扰以检验模型在干扰下的推理能力。结果表明,大多数先进VLMs(包括GPT-4)均易受干扰,推理能力明显下降。值得注意的是,InternVL2模型表现出更强的鲁棒性。此外,模型对文本干扰更为敏感。我们还探索了提示工程等缓解策略,虽能提升准确率,但仍有较大改进空间。

原文摘要 · Abstract (English)

Although vision-language models (VLMs) have achieved significant success in various applications such as visual question answering, their resilience to prompt variations remains an under-explored area. Understanding how distractions affect VLMs is crucial for improving their real-world applicability, as inputs could have noisy and irrelevant information in many practical scenarios. This paper aims to assess the robustness of VLMs against both visual and textual distractions in the context of science question answering. Built on the ScienceQA dataset, we developed a new benchmark that introduces distractions in both the visual and textual contexts to evaluate the reasoning capacity of VLMs amid these distractions. Our findings reveal that most-of-the-art VLMs, including GPT-4, are vulnerable to various types of distractions, experiencing noticeable degradation in reasoning capabilities when confronted with distractions. Notably, models such as InternVL2 demonstrate a higher degree of robustness to these distractions. We also found that models exhibit greater sensitivity to textual distractions than visual ones. Additionally, we explored various mitigation strategies, such as prompt engineering, to counteract the impact of distractions. While these strategies improved solution accuracy, our analysis shows that there remain significant opportunities for improvement.

视觉语言模型鲁棒性干扰测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。