用词联想测试评估并缓解大模型的文化偏见。
From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test
- 设计自适应词联想任务,探测模型跨文化认知对齐度。
- 现有模型在词联想中明显偏好西方(尤其是美国)认知模式。
- 提出CultureSteer,通过内嵌文化语义关联提升跨文化对齐。
人类中心的词联想测试(WAT)作为认知代理,通过共享的文化语义预期和由生活经验塑造的隐性语言模式,揭示社会文化差异。我们将该测试拓展为适用于大语言模型(LLMs)的自由关联任务,以评估其与跨文化认知的对齐程度。针对文化偏好问题,提出CultureSteer方法,不再依赖表面文化提示,而是将文化特定的语义关联直接嵌入模型内部表示空间。实验表明,当前大模型在词联想层面显著偏向西方(尤其美国)认知模式。相比之下,本模型显著提升跨文化对齐,捕捉多样化的语义关联。在文化敏感的下游任务上进一步验证了其有效性,证明其能促进跨文化认知对齐。本工作提出了增强大模型文化意识的新范式,推动更包容的语言技术发展。
原文摘要 · Abstract (English)
The human-centered word association test (WAT) serves as a cognitive proxy, revealing sociocultural variations through culturally shared semantic expectations and implicit linguistic patterns shaped by lived experiences. We extend this test into an LLM-adaptive, free-relation task to assess the alignment of large language models (LLMs) with cross-cultural cognition. To address culture preference, we propose CultureSteer, an innovative approach that moves beyond superficial cultural prompting by embedding cultural-specific semantic associations directly within the model's internal representation space. Experiments show that current LLMs exhibit significant bias toward Western (notably American) schemas at the word association level. In contrast, our model substantially improves cross-cultural alignment, capturing diverse semantic associations. Further validation on culture-sensitive downstream tasks confirms its efficacy in fostering cognitive alignment across cultures. This work contributes a novel methodological paradigm for enhancing cultural awareness in LLMs, advancing the development of more inclusive language technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。