arXiv:2511.14476cs.AI2025-11AAAI被引 13

研究不同群体价值观对大模型对齐的影响,发现安全与包容性存在权衡。

Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior

  • 按不同社会群体数据微调模型,测试价值观多样性对行为的影响
  • 男性比女性低18%毒害评分,黑人群体情感意识高44%
  • 保留评价分歧可使毒性降低53%,五级评分比二分制好22%

尽管大语言模型(LLMs)越来越多地通过人类反馈进行安全对齐,但对齐决策常忽视社会多样性。本研究系统评估了对齐流程中的人口统计差异与设计参数,收集了1,095名美国和德国参与者(共27,375次评分)对模型响应在毒性、情感意识(EA)、敏感性、刻板偏见和帮助性五个维度的评价。我们使用不同社会群体的偏好微调多个大语言模型和大推理模型,同时改变评分量表、分歧处理方法和优化技术。结果显示系统性人口差异:男性参与者比女性低18%毒害评分;保守派和黑人参与者的情感意识评分分别比自由派和白人高出27.9%和44%。基于群体特定偏好微调的模型表现出明显不同的行为。技术设计选择影响显著:保留评价分歧带来的毒性减少比多数投票高出约53%,五级量表比二分制带来约22%的改善;直接偏好优化(DPO)在多价值优化中始终优于组相对策略优化(GRPO)。这些发现为关键问题迈出初步一步:对齐应如何平衡专家信号与用户信号,以兼顾安全性与公平代表性?

原文摘要 · Abstract (English)

Although large language models (LLMs) are increasingly trained using human feedback for safety and alignment with human values, alignment decisions often overlook human social diversity. This study examines how incorporating pluralistic values affects LLM behavior by systematically evaluating demographic variation and design parameters in the alignment pipeline. We collect alignment data from US and German participants (N = 1,095 participants, 27,375 ratings) who rated LLM responses across five dimensions: Toxicity, Emotional Awareness (EA), Sensitivity, Stereotypical Bias, and Helpfulness. We fine-tuned multiple Large Language Models and Large Reasoning Models using preferences from different social groups while varying rating scales, disagreement handling methods, and optimization techniques. The results revealed systematic demographic effects: male participants rated responses 18% less toxic than female participants; conservative and Black participants rated responses 27.9% and 44% higher on EA than liberal and White participants, respectively. Models fine-tuned on group-specific preferences exhibited distinct behaviors. Technical design choices showed strong effects: the preservation of rater disagreement achieved roughly 53% greater toxicity reduction than majority voting, and 5-point scales yielded about 22% more reduction than binary formats; and Direct Preference Optimization (DPO) consistently outperformed Group Relative Policy Optimization (GRPO) in multi-value optimization. These findings represent a preliminary step in answering a critical question: How should alignment balance expert-driven and user-driven signals to ensure both safety and fair representation?

大模型对齐价值观多样性公平性人类反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。