arXiv:2605.26878cs.AI2026-05

解决多利益相关者模型评分中的权重不稳问题

Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation

论文配图:Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation
图 1 · 摘自论文原文
  • 分离评估与聚合,先固定权重再独立估算各角色效用
  • 实验显示权重噪声随利益方增多显著增大,最高达40%
  • 适合需要公平评分的多用户决策系统

多利益相关者任务要求单一输出满足具有冲突偏好的用户。现有全盘评估的LLM评判机制将效用估计与效用聚合混为一谈,导致隐式权重不稳定。我们通过理论与实证发现,这种聚合特有的‘权重噪声’在利益方满意度分散时会引起显著分数波动;实验中,该噪声随利益方数量增加而上升,最大增幅达40%。为此提出DecompR:基于查询结构预先固定反事实校准权重,候选输出评分前独立估计各角色效用,消除依赖候选的权重漂移,降低估计噪声。

原文摘要 · Abstract (English)

Multi-stakeholder tasks require one output to satisfy users with conflicting preferences. Holistic LLM judges conflate utility estimation and utility aggregation, yielding unstable implicit weights. We show empirically and theoretically that this aggregation-specific \emph{weighting noise} can create large score shifts when stakeholder satisfaction is dispersed; in our experiments, these weight-induced shifts also increase with stakeholder count. We propose \textsc{DecompR}: counterfactual-calibrated weights are fixed from query structure before candidate scoring, while per-role utilities are estimated independently, removing candidate-dependent weight drift and reducing estimation noise.

大模型对齐多主体评分系统权重稳定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。