测试多模态模型在视觉化选项下判断文化价值观的稳定性
ValueGround: Evaluating Culture-Conditioned Visual Value Grounding in MLLMs
- 用对比图像对模拟文化立场,替代文字选项进行评估
- 模型在视觉选项下的准确率从72.8%降至62.6%
- 适合研究跨模态文化认知与多模态大模型偏差
文化价值观不仅通过语言表达,也体现在视觉场景和日常社会实践中。然而,现有语言模型的文化价值观评估几乎完全依赖文本,无法判断当回答选项被可视化时,文化相关判断是否仍稳定。我们提出ValueGround,一个用于评估多模态大语言模型(MLLMs)中文化条件视觉价值定位的基准。该基准基于世界价值观调查(World Values Survey)的问题,使用最小对比的图像对表示对立的回答选项,同时控制无关变量。给定国家、问题和图像对,模型需在不接触原始文本选项的情况下,选择最符合该国价值倾向的图像。在六种MLLM和13个国家上的实验表明,模型在视觉选项下的表现显著下降,平均准确率从72.8%降至62.6%。本基准为研究文化条件价值判断的跨模态迁移提供了受控测试平台。
原文摘要 · Abstract (English)
Cultural values are expressed not only through language but also through visual scenes and everyday social practices. Yet existing evaluations of cultural values in language models are almost entirely text-only, leaving it unclear whether culture-conditioned judgments remain stable when response options are visualized. We introduce ValueGround, a benchmark for evaluating culture-conditioned visual value grounding in multimodal large language models (MLLMs). Built from World Values Survey questions, ValueGround uses minimally contrastive image pairs to represent opposing response options while controlling irrelevant variation. Given a country, a question, and an image pair, a model must choose the image that best matches the country's value tendency without access to the original response-option texts. Experiments across six MLLMs and 13 countries show that models perform substantially worse with visualized response options than with the original textual options, with average accuracy dropping from 72.8% to 62.6%. Our benchmark provides a controlled testbed for studying cross-modal transfer of culture-conditioned value judgments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。