发现大模型文化定位结果不可靠,噪声远超信号
Measurement Validity in LLM Cultural Alignment

- 拆解模型回答中的噪声、提示词偏差和随机性
- 42%的模型-问题组合信噪比超1.0,最高达5.56
- 提示语气变化可让模型文化位置偏移2.4单位
研究人员常将大模型的问卷回答视为人类文化价值观的代理。这包括将模型输出映射到如英格尔哈特-维尔策文化地图等工具,并推断模型所代表的文化类型。然而,模型对价值敏感问题的回答可能携带采样噪声,且对问题表述高度敏感。本文对多个大模型的调查回答进行分解,分离出随机种子、提示重述带来的变异。采用信噪比(NSR)检验模型看似文化位置是否显著区别于噪声。在涵盖四个地理来源的十余个模型、针对88个综合价值观调查国家的测试中,答案大多为否:117个有效模型-问题对中有49个(42%)的NSR超过1.0,最差情况达到5.56。两个模型甚至无法完成足够数量的问卷问题。结果印证了先前发现——大模型倾向于西方英语文化位置。但本研究揭示,当前尚无法精确解读特定模型的文化坐标:仅提示语气改变就可使模型位置偏移2.4个地图单位,相当于真实国家间的距离。这表明,基于大模型问卷回答进行文化归因前,必须先验证测量的可靠性。
原文摘要 · Abstract (English)
Researchers increasingly treat LLM survey responses as a proxy for human cultural values. This includes projecting model outputs onto instruments like the Inglehart-Welzel Cultural Map and drawing conclusions about which cultures a model resembles. While a model's answer to a value-laden questions may be interpreted as a cultural signal, it also carries sampling noise and, can be quite sensitive to question framing. In this paper, we separate survey responses, sampling noise and question framing for multiple LLMs. We decompose response variance from these models into variation across random seeds, prompt rewordings. We employ noise-to-signal ratio (NSR) to test whether a model's apparent cultural position is distinguishable from noise. When applied across a dozen models from four geographic origins, calibrated against 88 Integrated Values Survey countries, the answer is often no. NSR exceeds 1.0 on 49 of 117 valid model-question pairs (42%), reaching 5.56 in the worst case. Two models even refuse to answer sufficient number of survey questions outright. Our results corroborate previous findings that LLMs cluster toward Western, English-speaking cultural positions. However, what does not hold up in this study is the precision with which anyone can currently interpret a specific model's coordinates: prompt tone alone can shift a model by 2.4 map units, comparable to the distance between actual countries in the Inglehart-Welzel Cultural Map. These findings suggest that cultural attribution from LLM survey responses requires establishing the reliability of the underlying measurements before interpreting model coordinates as evidence of cultural representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。