arXiv:2502.19158cs.CLcs.AI2025-02EMNLP被引 9

提出多维度评估框架,揭示个性化偏好学习的性能与安全风险

When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning

  • 构建包含性能、公平性、适应性的多维度评估体系
  • 用户分歧大时方法间性能差异达36%,安全错配最高20%
  • 适合关注模型公平性与个性化的研究者与开发者

尽管基于人类反馈的强化学习(RLHF)广泛用于对齐大语言模型与人类偏好,但通常假设用户偏好一致,忽视了多样化的价值观与少数观点。个性化偏好学习虽能为个体用户定制偏好,但缺乏标准化评估方法。本文提出多维度评估框架,不仅衡量性能,还考察公平性、意外影响及在不同偏好分歧水平下的适应能力。通过在三个偏好数据集上对比八种个性化方法的大量实验,发现当用户强烈分歧时,方法间性能差异可达36%,且个性化可能引入高达20%的安全性错配。这些结果凸显了全面评估方法对发展更高效、更包容的偏好学习系统的重要性。

原文摘要 · Abstract (English)

While Reinforcement Learning from Human Feedback (RLHF) is widely used to align Large Language Models (LLMs) with human preferences, it typically assumes homogeneous preferences across users, overlooking diverse human values and minority viewpoints. Although personalized preference learning addresses this by tailoring separate preferences for individual users, the field lacks standardized methods to assess its effectiveness. We present a multi-faceted evaluation framework that measures not only performance but also fairness, unintended effects, and adaptability across varying levels of preference divergence. Through extensive experiments comparing eight personalization methods across three preference datasets, we demonstrate that performance differences between methods could reach 36% when users strongly disagree, and personalization can introduce up to 20% safety misalignment. These findings highlight the critical need for holistic evaluation approaches to advance the development of more effective and inclusive preference learning systems.

偏好学习RLHF公平性LLM评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。