用微调方法改善大模型跨文化偏见,发现中国模型在本国人群表现最差。
Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning

- 针对最极端的5类人群做轻量微调,仅需1200条数据和15分钟训练。
- 波兰模型偏见降低16.8%,统计显著且所有目标人群均改善。
- 微调只是转移偏见而非消除,可能引发新的文化误判。
我们评估了三个开源大模型(美国Gemma3-12B、波兰Bielik-11B-v3、中国Qwen3-4B)在世界价值观调查第七波数据下,针对63个不同人口画像在三个国家中的分布偏差,采用归一化Wasserstein距离量化文化错位程度。出人意料的是,没有模型偏好其母国;其中由中国构建的Qwen3-4B在本国人群中的错位值最高(W1 = 0.436),为整个模型×国家矩阵中最大。对五个最差表现的人群进行基于LoRA的定向微调,仅需少于1,200个训练样本,单卡训练时间不足15分钟,使Bielik-11B的偏见减少16.8%(p_Bonf = 0.002,d = -4.4),且五个目标全部改善。然而,国家层面分解显示,微调并未消除偏见,而是将其重新分配:Bielik模型中表现最差的群体从美国老人转变为中国人老人,前后无重叠。据我们所知,这是首个针对最坏情况人口画像实施LoRA微调以缓解跨文化偏见的研究。
原文摘要 · Abstract (English)
We evaluate three open-weight LLMs (Gemma3-12B from the USA, Bielik-11B-v3 from Poland, and Qwen3-4B from China) against World Values Survey Wave 7 data for 63 demographic personas across three countries, using normalized Wasserstein distance to quantify distributional misalignment. Contrary to expectations, no model favors its home country: the Chinese-built Qwen3-4B performs worst on its own Chinese population (W1 = 0.436, the highest misalignment in the entire model x country matrix). Targeted LoRA fine-tuning on the five worst-case personas, requiring fewer than 1,200 training pairs and under 15 minutes on a single GPU, reduces bias by 16.8% for Bielik-11B (p_Bonf = 0.002, d = -4.4) with all five targets improving. However, country-level decomposition reveals that fine-tuning redistributes rather than removes bias: Bielik's worst-case personas swap entirely from American to Chinese elderly, with zero overlap between pre- and post-correction sets. To our knowledge, this is the first study to target worst-case demographic personas with LoRA fine-tuning for cross-cultural bias mitigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。