用多语言批判数据提升大模型文化敏感度,减少偏见。
CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis
- 通过合成多语言文化问题生成批判数据,增强模型文化理解。
- 在三个基准和新数据集上,模型文化对齐能力达开源最佳水平。
- 适合关注AI公平性、跨文化应用的研究者与开发者。
大语言模型在多项任务中表现出色,但常存在文化偏见,忽视低资源地区的价值观与语言多样性。这不仅损害普遍平等,还可能强化刻板印象与歧视。为此,我们提出CulFiT——一种基于多语言批判数据合成的文化感知训练范式。该方法合成多样文化相关问题,在文化相关语言中构建批判数据,并采用细粒度奖励机制,将文化文本分解为可验证的知识单元以实现可解释评估。我们还提出了GlobalCultureQA,一个用于评估全球语境下文化敏感回答的多语言开放问答数据集。在三个现有基准和我们的GlobalCultureQA上的大量实验表明,CulFiT在文化对齐和通用推理方面达到开源模型的最先进水平。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, yet they often exhibit a specific cultural biases, neglecting the values and linguistic diversity of low-resource regions. This cultural bias not only undermines universal equality, but also risks reinforcing stereotypes and perpetuating discrimination. To address this, we propose CulFiT, a novel culturally-aware training paradigm that leverages multilingual data and fine-grained reward modeling to enhance cultural sensitivity and inclusivity. Our approach synthesizes diverse cultural-related questions, constructs critique data in culturally relevant languages, and employs fine-grained rewards to decompose cultural texts into verifiable knowledge units for interpretable evaluation. We also introduce GlobalCultureQA, a multilingual open-ended question-answering dataset designed to evaluate culturally-aware responses in a global context. Extensive experiments on three existing benchmarks and our GlobalCultureQA demonstrate that CulFiT achieves state-of-the-art open-source model performance in cultural alignment and general reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。