arXiv:2502.14906cs.CLcs.AI2025-02NAACL被引 17

探究多模态模型如何受文化背景影响,发现图像能辅助理解文化价值但效果依赖具体场景。

Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models

  • 通过图文结合数据评估不同规模模型对文化价值的敏感性
  • 模型在文化价值对齐上表现不一,依赖具体使用上下文
  • 适合关注跨文化AI伦理与多模态对齐的研究者阅读

基于文化背景的大型语言模型(LLM)价值对齐研究已成为重要方向。然而,大规模视觉-语言模型(VLM)中的类似偏见尚未得到充分探索。随着多模态模型规模持续扩大,评估图像能否作为文化可靠代理,并考察视觉与文本数据融合后价值如何嵌入变得愈发关键。本文系统评估了不同规模的多模态模型在文化价值对齐方面的表现。结果表明,与LLM类似,VLM对文化价值表现出敏感性,但其对齐能力高度依赖具体上下文。尽管图像有助于提升对文化价值的理解,但该对齐效果在不同场景间差异显著,凸显多模态模型对齐中的复杂性与未充分探索的挑战。

原文摘要 · Abstract (English)

Investigating value alignment in Large Language Models (LLMs) based on cultural context has become a critical area of research. However, similar biases have not been extensively explored in large vision-language models (VLMs). As the scale of multimodal models continues to grow, it becomes increasingly important to assess whether images can serve as reliable proxies for culture and how these values are embedded through the integration of both visual and textual data. In this paper, we conduct a thorough evaluation of multimodal model at different scales, focusing on their alignment with cultural values. Our findings reveal that, much like LLMs, VLMs exhibit sensitivity to cultural values, but their performance in aligning with these values is highly context-dependent. While VLMs show potential in improving value understanding through the use of images, this alignment varies significantly across contexts highlighting the complexities and underexplored challenges in the alignment of multimodal models.

多模态对齐文化敏感性视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。