让奖励模型兼顾多元文化偏好,减少偏见。
Steerable Cultural Preference Optimization of Reward Models

- 设计新算法SCPO,动态平衡不同文化群体偏好。
- 在7国数据上,少数群体模型性能提升最高7分。
- 训练效率提升280%,适合多文化场景的对齐研究。
大语言模型要服务于不同文化子群体,但现有对齐研究多聚焦于特定地区标注者的统一偏好。本文旨在推动更具全球视野的对齐模型发展,使其准确反映子群体偏好且不偏向任一文化。聚焦奖励模型,提出新型训练算法SCPO,可均衡融合多元文化偏好。实验显示,在PRISM和GlobalOpinionQA两个数据集、7个国家中,少数群体奖励模型性能较基线最高提升7分;与全数据微调相比,训练数据效率最高提升280%。通过分别评估子群体偏好,验证了所提加权方法有效缓解过度偏见。代码已开源。
原文摘要 · Abstract (English)
It is essential for large language model (LLM) technology to serve many different cultural sub-communities in a manner that is acceptable to each community. However, research on LLM alignment has so far predominantly focused on predicting a unified response preference of annotators from certain regions. This paper aims to advance the development of alignment models with a more global outlook, that are able to accurately represent the preferences of subcommunities and do not exhibit excessive bias towards any of them. We focus on the development of reward models for this purpose and present a novel reward model training algorithm (SCPO) that can incorporate diverse cultural preferences in a balanced manner. Our method results in performance increases of the minority reward model of up to 7 points over the baseline model across two datasets, PRISM and GlobalOpinionQA, and across 7 countries. SCPO is up to 280% more training data-efficient than full-data finetuning of reward models. In addition, we perform analysis of bias by separately evaluating on the preference of subcommunities and show that excessive bias is mitigated via our weighting method. Our code is available at https://github.com/minsik-ai/Steerable-Cultural-Preference
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。