arXiv:2507.00700cs.CL2025-07

日语和英语训练的视觉语言模型展现不同认知风格。

Contrasting Cognitive Styles in Vision-Language Models: Holistic Attention in Japanese Versus Analytical Focus in English

  • 对比日英语料训练的模型注意力模式。
  • 日语模型更关注整体上下文,英语模型聚焦单个物体。
  • 适合研究文化影响与AI行为的学者参考。

跨文化感知与认知研究表明,不同文化背景的人处理视觉信息的方式存在差异。例如,东亚人倾向于整体性视角,关注上下文关系;而西方人则常采用分析性方法,聚焦于个体对象及其属性。本研究探讨以日语和英语为主训练的视觉语言模型(VLMs)是否表现出类似的文化根植注意力模式。通过比较图像描述的生成结果,我们发现这些模型不仅内化了语言的结构特征,还再现了训练数据中嵌入的文化行为,表明文化认知可能在隐式层面塑造模型输出。

原文摘要 · Abstract (English)

Cross-cultural research in perception and cognition has shown that individuals from different cultural backgrounds process visual information in distinct ways. East Asians, for example, tend to adopt a holistic perspective, attending to contextual relationships, whereas Westerners often employ an analytical approach, focusing on individual objects and their attributes. In this study, we investigate whether Vision-Language Models (VLMs) trained predominantly on different languages, specifically Japanese and English, exhibit similar culturally grounded attentional patterns. Using comparative analysis of image descriptions, we examine whether these models reflect differences in holistic versus analytic tendencies. Our findings suggest that VLMs not only internalize the structural properties of language but also reproduce cultural behaviors embedded in the training data, indicating that cultural cognition may implicitly shape model outputs.

视觉语言模型文化认知注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。