arXiv:2509.21798cs.CLcs.AI2025-09被引 2

构建文化感知评估基准,提升大模型对多元文化的理解与对齐能力

Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment

  • 提出跨10种文化、4个领域的评估基准CARB,填补文化评估数据空白
  • 发现现有奖励模型在文化理解上表现不足,且评分易依赖表面特征
  • 引入‘本地思维’机制,通过可验证奖励强化学习,提升文化推理深度

奖励模型(RMs)在使大语言模型(LLMs)与多元文化对齐中至关重要。然而,现有评估因缺乏文化相关数据集而难以衡量文化意识。为此,我们提出了文化感知奖励建模基准(CARB),覆盖10种不同文化及4个文化领域。对前沿奖励模型的广泛评估揭示其在文化建模上的不足,并证明在CARB上的表现与下游多语言文化对齐任务正相关。进一步分析发现,文化感知奖励建模中存在虚假相关性,即模型评分主要依赖表面特征而非真实文化内涵理解。为此,我们提出Think-as-Locals方法,通过可验证奖励强化学习(RLVR)引导生成式奖励模型进行更深层的文化化推理,并设计合理奖励以确保偏好判断准确及高质量结构化评估标准生成。实验验证该方法有效缓解虚假特征干扰,推动文化感知奖励建模发展。

原文摘要 · Abstract (English)

Reward models (RMs) are crucial for aligning large language models (LLMs) with diverse cultures. Consequently, evaluating their cultural awareness is essential for further advancing global alignment of LLMs. However, existing RM evaluations fall short in assessing cultural awareness due to the scarcity of culturally relevant evaluation datasets. To fill this gap, we propose Cultural Awareness Reward modeling Benchmark (CARB), covering 10 distinct cultures across 4 cultural domains. Our extensive evaluation of state-of-the-art RMs reveals their deficiencies in modeling cultural awareness and demonstrates a positive correlation between performance on CARB and downstream multilingual cultural alignment tasks. Further analysis identifies the spurious correlations within culture-aware reward modeling, wherein RM's scoring relies predominantly on surface-level features rather than authentic cultural nuance understanding. To address these, we propose Think-as-Locals to elicit deeper culturally grounded reasoning from generative RMs via reinforcement learning from verifiable rewards (RLVR) and employ well-designed rewards to ensure accurate preference judgments and high-quality structured evaluation criteria generation. Experimental results validate its efficacy in mitigating spurious features interference and advancing culture-aware reward modeling.

奖励模型文化对齐评估基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。