测试主流大模型的文化倾向,发现多数默认偏美,文化提示可有效调整。
The American Ghost in the Machine: How language models align culturally and the effects of cultural prompting
- 用文化维度理论分析8个大模型的文化偏好
- 7个模型在文化提示下能转向目标国家文化
- 中日文化对齐困难,即使中文模型也难突破
文化是人类互动的基础,决定我们如何感知和回应日常交流。随着生成式大语言模型(LLMs)在人机交互中的普及,模型的文化对齐成为重要研究课题。本文基于VSM13国际调查与霍夫斯泰德文化维度理论,分析了DeepSeek-V3、V3.1、GPT-5、GPT-4.1、GPT-4、Claude Opus 4、Llama 3.1和Mistral Large等8个主流大模型的文化倾向。随后通过文化提示(使用系统提示将模型对齐至特定国家),测试其向中国、法国、印度、伊朗、日本和美国等文化的适应能力。结果显示,在未指定文化时,多数模型默认倾向美国;在文化提示下,7个模型能向目标文化靠拢。但模型在对齐日本与中国文化时表现不佳,即便其中两个模型由中文公司深度求索开发。
原文摘要 · Abstract (English)
Culture is the bedrock of human interaction; it dictates how we perceive and respond to everyday interactions. As the field of human-computer interaction grows via the rise of generative Large Language Models (LLMs), the cultural alignment of these models become an important field of study. This work, using the VSM13 International Survey and Hofstede's cultural dimensions, identifies the cultural alignment of popular LLMs (DeepSeek-V3, V3.1, GPT-5, GPT-4.1, GPT-4, Claude Opus 4, Llama 3.1, and Mistral Large). We then use cultural prompting, or using system prompts to shift the cultural alignment of a model to a desired country, to test the adaptability of these models to other cultures, namely China, France, India, Iran, Japan, and the United States. We find that the majority of the eight LLMs tested favor the United States when the culture is not specified, with varying results when prompted for other cultures. When using cultural prompting, seven of the eight models shifted closer to the expected culture. We find that models had trouble aligning with Japan and China, despite two of the models tested originating with the Chinese company DeepSeek.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。