arXiv:2510.23965cs.AIcs.LG2025-10被引 2

解决大模型对齐中用户偏好差异问题,提升平均偏好估计准确性。

The Sign Estimator: LLM Alignment in the Face of Choice Heterogeneity

  • 用二分类损失替代交叉熵,实现稳定有序对齐
  • 在模拟中误差降低35%,与真实偏好一致率从12%提至8%
  • 无需追踪个体数据,兼容现有对齐流程

传统大模型对齐方法易受人类偏好异质性影响。对成对比较数据(如提示-完成对)拟合朴素概率模型会得到不一致的群体平均效用估计——这是社会福利的标准度量。本文提出一种新方法,称为符号估计器(sign estimator),通过在聚合步骤中以二分类损失替换交叉熵,提供简单、可证明一致且高效的估计。该修改在弱假设下恢复了稳定的序数对齐,并首次在此设定下获得多项式阶有限样本误差界。在使用数字孪生进行的大模型对齐真实模拟中,符号估计器显著降低了偏好扭曲,在一组模拟人物中将(角度)估计误差降低近35%,并将与真实群体偏好的分歧从12%降至8%,优于标准RLHF。该方法还优于需追踪个体偏好数据的面板数据启发式方法,同时保持现有对齐流水线的实现简洁性。

原文摘要 · Abstract (English)

Traditional LLM alignment methods are vulnerable to heterogeneity in human preferences. Fitting a naïve probabilistic model to pairwise comparison data (say over prompt-completion pairs) yields an inconsistent estimate of the population-average utility -a canonical measure of social welfare. We propose a new method, dubbed the sign estimator, that provides a simple, provably consistent, and efficient estimator by replacing cross-entropy with binary classification loss in the aggregation step. This simple modification recovers consistent ordinal alignment under mild assumptions and achieves the first polynomial finite-sample error bounds in this setting. In realistic simulations of LLM alignment using digital twins, the sign estimator substantially reduces preference distortion over a panel of simulated personas, cutting (angular) estimation error by nearly 35% and decreasing disagreement with true population preferences from 12% to 8% compared to standard RLHF. Our method also compares favorably to panel data heuristics that explicitly model user heterogeneity and require tracking individual-level preference data-all while maintaining the implementation simplicity of existing LLM alignment pipelines.

大模型对齐偏好建模效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。