arXiv:2511.22787cs.CV2025-11被引 2

研究视觉语言模型在文化混杂场景下的表现,发现其易受背景干扰且识别不一致。

World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models

  • 构建2.3万张混合文化图像的VQA数据集,系统测试模型表现。
  • 模型在文化混杂时准确率下降14%,对相同食物的判断因上下文不一致。
  • 通过多样化数据微调可显著提升模型鲁棒性,适合关注真实场景应用的研究者。

在全球化背景下,不同来源的文化元素常共现于同一视觉场景中,我们称之为文化混杂场景。现有大视觉语言模型(LVLMs)对此类场景的理解仍不充分。本文将文化混杂视为对LVLMs的关键挑战,系统考察多文化元素共现时模型的行为表现。为此,我们构建了CultureMix——一个包含23,000张扩散生成、人工验证的文化混杂图像的食品视觉问答(VQA)基准,涵盖四个子任务:(1) 仅食物,(2) 食物+食物,(3) 食物+背景,(4) 食物+食物+背景。评估10个主流LVLM后发现,模型在混杂环境中普遍存在文化身份丢失问题,表现出强烈背景依赖性:相比仅食物基线,添加文化背景使准确率下降14%;对相同食物在不同语境下产生不一致预测。为缓解此问题,我们探索三种鲁棒性策略,发现使用多样化文化混杂数据进行监督微调能显著提升模型一致性并降低对背景的敏感性。本文呼吁重视文化混杂场景,推动模型在多元文化真实环境中的可靠部署。

原文摘要 · Abstract (English)

In a globalized world, cultural elements from diverse origins frequently appear together within a single visual scene. We refer to these as culture mixing scenarios, yet how Large Vision-Language Models (LVLMs) perceive them remains underexplored. We investigate culture mixing as a critical challenge for LVLMs and examine how current models behave when cultural items from multiple regions appear together. To systematically analyze these behaviors, we construct CultureMix, a food Visual Question Answering (VQA) benchmark with 23k diffusion-generated, human-verified culture mixing images across four subtasks: (1) food-only, (2) food+food, (3) food+background, and (4) food+food+background. Evaluating 10 LVLMs, we find consistent failures to preserve individual cultural identities in mixed settings. Models show strong background reliance, with accuracy dropping 14% when cultural backgrounds are added to food-only baselines, and they produce inconsistent predictions for identical foods across different contexts. To address these limitations, we explore three robustness strategies. We find supervised fine-tuning using a diverse culture mixing dataset substantially improve model consistency and reduce background sensitivity. We call for increased attention to culture mixing scenarios as a critical step toward developing LVLMs capable of operating reliably in culturally diverse real-world environments.

视觉语言模型文化混杂鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。