arXiv:2504.05154cs.CL2025-04EMNLP被引 16

用3500个文化特异性问题提升大模型跨文化理解能力

CARE: Multilingual Human Preference Learning for Cultural Awareness

  • 构建多语言文化偏好数据集CARE,含3490个问题和31.7万条人类评分回复
  • 少量高质量本地化偏好数据可显著提升模型跨文化表现,优于大规模通用数据
  • 适合关注跨文化对齐、多语言AI公平性的研究者与开发者

语言模型通常通过人类偏好进行调优以生成有用回复,但偏好调优对处理文化多样性请求的影响仍缺乏研究。本文系统分析如何将母语者的文化偏好融入偏好学习过程,以训练更具文化意识的语言模型。我们提出 extbf{CARE},一个包含3,490个文化特异性问题和31.7万条带人类评分的回复的多语言资源。实验表明,少量高质量的本地化偏好数据能显著提升多种语言模型的文化敏感性,优于更大规模的通用偏好数据。分析显示,初始文化表现较强的模型在对齐后收益更明显,导致不同地区模型因数据获取差异产生性能差距。CARE已开源:https://github.com/Guochry/CARE。

原文摘要 · Abstract (English)

Language Models (LMs) are typically tuned with human preferences to produce helpful responses, but the impact of preference tuning on the ability to handle culturally diverse queries remains understudied. In this paper, we systematically analyze how native human cultural preferences can be incorporated into the preference learning process to train more culturally aware LMs. We introduce \textbf{CARE}, a multilingual resource containing 3,490 culturally specific questions and 31.7k responses with human judgments. We demonstrate how a modest amount of high-quality native preferences improves cultural awareness across various LMs, outperforming larger generic preference data. Our analyses reveal that models with stronger initial cultural performance benefit more from alignment, leading to gaps among models developed in different regions with varying access to culturally relevant data. CARE is publicly available at https://github.com/Guochry/CARE.

文化感知多语言偏好学习数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。