用母语者词汇联想数据微调大模型,提升文化对齐效果。
ALIGN: Word Association Learning for Cultural Alignment in Large Language Models
- 基于母语者词汇联想数据进行微调,结合认知心理学机制。
- 中文模型精准度提升165%,美中价值观差异题响应对齐翻倍。
- 7-8B小模型媲美70B基线,适合资源有限的场景应用。
大型语言模型因训练数据中观点过拟而存在文化偏见,但文化对齐仍面临知识不足与有效学习方法缺乏的挑战。本文提出一种低成本、基于认知心理学的方法:在母语者词汇联想数据上微调大模型,利用联想关系捕捉文化知识。使用美国(英语)和中国(普通话)母语者的词联数据,对 Llama-3.1-8B 与 Qwen-2.5-7B 进行监督微调与偏好优化。通过两层评估框架检验词汇关联与文化价值对齐,采用世界价值观调查数据。结果表明,词汇对齐显著提升(英文精确率@5提高16-20%,中文提升43-165%),高阶文化价值观亦发生明显转变。在中美受访者分歧最大的50个问题上,微调后的Qwen模型对中文价值观的响应对齐率从13%升至25%。令人惊讶的是,7-8B模型表现媲美甚至超越未微调的70B基线,证明数百万条文化相关的词联数据即可实现价值对齐,无需昂贵重训练。本工作凸显了基于人类认知的研究在提升模型文化对齐中的潜力与必要性。
原文摘要 · Abstract (English)
Large language models (LLMs) exhibit cultural bias from overrepresented viewpoints in training data, yet cultural alignment remains a challenge due to limited cultural knowledge and a lack of exploration into effective learning approaches. We introduce a cost-efficient and cognitively grounded method: fine-tuning LLMs on native speakers' word-association norms, leveraging cognitive psychology findings that such associations capture cultural knowledge. Using word association datasets from native speakers in the US (English) and China (Mandarin), we train Llama-3.1-8B and Qwen-2.5-7B via supervised fine-tuning and preference optimization. We evaluate models' cultural alignment through a two-tier evaluation framework that spans lexical associations and cultural value alignment using the World Values Survey. Results show significant improvements in lexical alignment (16-20% English, 43-165% Mandarin on Precision@5) and high-level cultural value shifts. On a subset of 50 questions where US and Chinese respondents diverge most, fine-tuned Qwen nearly doubles its response alignment with Chinese values (13 to 25). Remarkably, our trained 7-8B models match or exceed vanilla 70B baselines, demonstrating that a few million of culture-grounded associations achieve value alignment without expensive retraining. Our work highlights both the promise and the need for future research grounded in human cognition in improving cultural alignment in AI models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。