通过优化数据代表性与独特性,提升大模型对15种文化的精准适配。
CAReDiO: Cultural Alignment via Representativeness and Distinctiveness Guided Data Optimization
- 基于信息论目标,交替优化文化相关问答数据
- 仅用200样本即实现小模型高效文化对齐
- 适合需跨文化适配的开源或商用大模型
随着大语言模型在多元地区部署,使其与多样文化对齐对提升用户参与度、减少文化冲突至关重要。现有方法依赖人工标注或合成的文化特异性语料库,但存在两大问题:(1) 代表性不足,无法充分覆盖目标文化核心特征,导致信息冗余;(2) 独特性缺失,难以区分目标文化与相关文化间的独特差异,影响精准建模。为此,我们提出CAReDiO,一种新型数据优化框架,通过在上下文学习中交替优化文化敏感型问答,依据两个信息论目标,同时增强数据的文化代表性与独特性。在15种文化上的大量实验表明,CAReDiO可生成富含文化信息的高质量数据,仅用200个训练样本即可实现小型开源模型或大型专有模型的有效文化对齐,在多项选择与开放问答基准上均优于先前数据集。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) are deployed across diverse regions, aligning them with pluralistic cultures is crucial for improving user engagement and mitigating cultural conflicts. Recent work has curated, either synthesized or manually annotated, culture-specific corpora for alignment. Nevertheless, inspired by cultural theories, we recognize they face two key challenges. (1) Representativeness: These corpora inadequately capture the target culture's core characteristics, causing insufficient cultural coverage and redundancy; (2) Distinctiveness: They fail to distinguish the unique nuances of the target culture from patterns shared across relevant ones, hindering precise culture modeling. To handle these challenges, we introduce CAReDiO, a novel data optimization framework that alternately optimizes culture-sensitive questions and responses according to two information-theoretic objectives in an in-context manner, enhancing both cultural representativeness and distinctiveness of constructed data. Extensive experiments on 15 cultures demonstrate that CAReDiO can create high-quality data with richer cultural information and enable efficient alignment of small open-source or large proprietary LLMs with as few as 200 training samples, consistently outperforming previous datasets in both multi-choice and open-ended benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。