arXiv:2509.08022cs.CLcs.AI2025-09中稿 · IJCAI

构建跨74国价值观对齐评测基准,揭示模型偏见并提升公平性

DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values

  • 基于74国用户反馈构建多维度价值观评测数据集
  • 发现模型在不同地区/人群间存在显著价值偏差
  • 轻量微调可同时提升本地与跨域对齐效果

将大语言模型与多元人类价值观对齐对于安全有效部署至关重要,但现有评测常忽略文化与人口差异。我们提出DiverValue-Bench,一个面向74个国家/地区的群体感知评测基准。该基准包含23,763条经质量控制的实例,源自PRISM用户反馈,并通过大规模人工验证审核,具备细粒度价值标签、个性化问题、对比参考答案及丰富人口统计元数据。利用DiverValue-Bench,我们评估了代表性LLMs,发现聚合性能掩盖了显著的地理与人口差异。进一步表明,采用低秩适配(LoRA)和直接偏好优化(DPO)的轻量级偏好微调,能显著提升领域内对齐效果,并带来一致的域外增益。结果凸显了群体感知对齐评估的重要性,证明DiverValue-Bench是实现全球对齐、个性化价值建模与公平AI发展的实用基础。

原文摘要 · Abstract (English)

Aligning large language models (LLMs) with diverse human values is essential for safe and effective deployment, yet existing benchmarks often overlook cultural and demographic variation. We introduce DiverValue-Bench, a population-aware benchmark for evaluating multi-dimensional value alignment across 74 countries/regions. It contains 23,763 quality-controlled instances derived from PRISM user feedback and audited through large-scale human validation, with fine-grained value labels, personalized questions, contrastive reference answers, and rich demographic metadata. Using DiverValue-Bench, we evaluate representative LLMs and reveal substantial geographic and demographic disparities that are masked by aggregate performance. We further show that lightweight preference-based fine-tuning with Low-Rank Adaptation (LoRA) and Direct Preference Optimization (DPO) substantially improves in-domain value alignment while yielding consistent out-of-domain gains. These results highlight the need for population-aware alignment evaluation and demonstrate the utility of DiverValue-Bench as a practical foundation for global alignment, personalized value modeling, and equitable AI development.

价值观对齐多国评测公平AI轻量微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。