测试视觉语言模型在文化混杂场景下的稳定性,发现其易被误导
Vision Language Models are Confused Tourists
- 设计新评测集ConfusedTourist,模拟多重文化线索干扰
- 图像叠加后准确率显著下降,生成式扰动更致命
- 模型注意力被无关文化符号干扰,适合关注多文化应用的研究者
尽管文化维度是评估视觉-语言模型(VLMs)的关键方面,但其在多样文化输入下的稳定性仍缺乏充分检验,而这对于支持多元文化社会至关重要。现有评估通常每张图像仅包含单一文化概念,忽略了多个可能无关的文化线索共存的场景。为此,我们提出ConfusedTourist——一个新型文化对抗鲁棒性评测套件,用于评估VLMs在地理线索扰动下的稳定性。实验显示,简单图像堆叠扰动即导致准确率大幅下降,且基于图像生成的扰动使性能进一步恶化。可解释性分析表明,这些失败源于模型注意力系统性地转向干扰性线索,偏离了预期关注点。研究揭示:视觉文化概念混合会显著削弱甚至最先进的VLMs,凸显了构建更具文化鲁棒性的多模态理解系统的紧迫性。
原文摘要 · Abstract (English)
Although the cultural dimension has been one of the key aspects in evaluating Vision-Language Models (VLMs), their ability to remain stable across diverse cultural inputs remains largely untested, despite being crucial to support diversity and multicultural societies. Existing evaluations often rely on benchmarks featuring only a singular cultural concept per image, overlooking scenarios where multiple, potentially unrelated cultural cues coexist. To address this gap, we introduce ConfusedTourist, a novel cultural adversarial robustness suite designed to assess VLMs' stability against perturbed geographical cues. Our experiments reveal a critical vulnerability, where accuracy drops heavily under simple image-stacking perturbations and even worsens with its image-generation-based variant. Interpretability analyses further show that these failures stem from systematic attention shifts toward distracting cues, diverting the model from its intended focus. These findings highlight a critical challenge: visual cultural concept mixing can substantially impair even state-of-the-art VLMs, underscoring the urgent need for more culturally robust multimodal understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。