首个评估大模型视觉误导鲁棒性的综合基准,发现模型易受视觉干扰影响。
MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs
- 按视觉概念、属性、关系分层设计误导输入,构建六类1248个问答样本。
- 测试18个顶尖模型均暴露严重脆弱性,证明现有模型对视觉干扰不敏感。
- 提供细粒度鲁棒性指标MVI-Sensitivity,适合模型开发者与评测人员使用。
评估大视觉语言模型(LVLMs)的鲁棒性对其实用化部署至关重要。然而,现有评测基准多关注幻觉或误导性文本输入,忽视了视觉输入带来的同等重要挑战。为此,我们提出MVI-Bench,首个专为评估误导性视觉输入对LVLM鲁棒性影响而设计的综合性基准。基于基本视觉原语,其设计涵盖视觉概念、属性和关系三个层次,共构建六类代表性类别,并整理出1,248个专家标注的视觉问答实例。为支持细粒度评估,我们引入MVI-Sensitivity这一新指标,可从微观层面刻画模型鲁棒性。在18个先进LVLM上的实证结果揭示其对误导性视觉输入存在显著脆弱性。深入分析为提升模型可靠性提供了可操作的洞见。基准与代码库已开源:https://github.com/chenyil6/MVI-Bench。
原文摘要 · Abstract (English)
Evaluating the robustness of Large Vision-Language Models (LVLMs) is essential for their continued development and responsible deployment in real-world applications. However, existing robustness benchmarks typically focus on hallucination or misleading textual inputs, while largely overlooking the equally critical challenge posed by misleading visual inputs in assessing visual understanding. To fill this important gap, we introduce MVI-Bench, the first comprehensive benchmark specially designed for evaluating how Misleading Visual Inputs undermine the robustness of LVLMs. Grounded in fundamental visual primitives, the design of MVI-Bench centers on three hierarchical levels of misleading visual inputs: Visual Concept, Visual Attribute, and Visual Relationship. Using this taxonomy, we curate six representative categories and compile 1,248 expertly annotated VQA instances. To facilitate fine-grained robustness evaluation, we further introduce MVI-Sensitivity, a novel metric that characterizes LVLM robustness at a granular level. Empirical results across 18 state-of-the-art LVLMs uncover pronounced vulnerabilities to misleading visual inputs, and our in-depth analyses on MVI-Bench provide actionable insights that can guide the development of more reliable and robust LVLMs. The benchmark and codebase can be accessed at https://github.com/chenyil6/MVI-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。