让大模型根据文化语境选择合适实体,避免无差别偏见。
CoCoA: Context-Conditional Cultural Alignment for Large Language Models

- 通过双语境训练学习文化线索下的实体选择策略
- 文化偏见得分从43降至24,保持近中性偏好(50.2)
- 适合关注多语言文化适配的AI研发人员
大型语言模型常在各类文化背景下偏向西方相关实体。传统去偏方法追求统一中立,但文化偏见缓解需要依据语境调整行为:有文化线索时优选符合文化的实体,无线索时保持中立。我们提出CoCoA(上下文条件文化对齐)框架,通过在同一实体对上于有/无文化线索的双语境下进行训练,结合对比对齐目标、校准与漂移正则化,并利用目标感知梯度协调优化。我们在CAMeL和Camellia两个以实体为中心的文化偏见基准上,针对十种语言设置和四种LLM进行了评估。CoCoA将平均文化偏见得分从43降至24,同时保持中性偏好为50.2,对五个标准基准的通用性能影响极小。结果表明,有效文化对齐需依赖上下文条件建模而非统一去偏,为缓解LLM中的实体中心文化偏见开辟新方向。
原文摘要 · Abstract (English)
Large Language Models (LLMs) often favor Western-associated entities across cultural contexts. Conventional debiasing methods aim for uniform neutrality, but cultural bias mitigation demands context-conditional behavior, preferring culturally appropriate entities when cultural cues are present and remaining neutral when they are absent. We propose CoCoA (Context-Conditional Cultural Alignment), a framework that learns this behavior through dual-context training on the same entity pairs under contexts with and without cultural cues. CoCoA combines a contrastive alignment objective with calibration and drift regularization, optimized through goal-aware gradient reconciliation. We evaluate CoCoA on CAMeL and Camellia, two entity-centric cultural bias benchmarks, across ten language settings and four LLMs. CoCoA reduces the Cultural Bias Score from 43 to 24 on average while maintaining near-neutral preferences at 50.2, with minimal impact on general performance across five standard benchmarks. These findings highlight that effective cultural alignment requires context-conditional modeling rather than uniform debiasing, and establish a new direction for mitigating entity-centric cultural bias in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。