arXiv:2607.20862cs.CL2026-07

通过融合多个专家模型的隐藏状态,提升不可验证偏好任务的评估效果。

CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation

论文配图:CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation
图 1 · 摘自论文原文
  • 将多个冻结的奖励模型视为互补评估者,融合其隐藏状态
  • 在两个基准数据集上优于单模型和多模型评分基线
  • 适合需要整合多种评估视角的复杂偏好判断场景

目前,对不可验证任务的可靠评估仍具挑战性。现有方法往往难以充分捕捉人类偏好背后的多样化评价标准。为此,我们提出约束性共享-私有融合(CSPF)方法,将异构的冻结奖励模型视为互补评估者,在成对人类偏好监督下学习融合其隐藏状态表示。CSPF将每个专家信号分解为共享与专家私有表示,促进跨专家对齐同时保留互补观点。在LM-Arena目标域适应和PPE分布外偏好评估实验中,CSPF在主要指标上优于单专家奖励模型、标量得分多专家及评分规则基线。总体表明,融合隐藏状态为偏好评估提供了更具表现力的基础,为不可验证偏好任务的集成评估信号提供了一条实用路径。

原文摘要 · Abstract (English)

At present, reliable evaluation of non-verifiable tasks remains challenging. Existing approaches often fail to adequately capture the diverse evaluative criteria underlying human preferences in such tasks. To this end, we propose Constrained Shared-Private Fusion (CSPF), a fusion method that treats heterogeneous frozen reward models as complementary evaluators and learns to integrate their hidden-state representations under pairwise human-preference supervision. CSPF decomposes each expert signal into shared and expert-private representations, encouraging cross-expert alignment while preserving complementary viewpoints. Across experiments on LM-Arena target-domain adaptation and PPE out-of-distribution preference evaluation, CSPF achieves the best performance on the primary metrics among the evaluated single-expert reward-model, scalar-score multi-expert, and rubric-judge baselines. Overall, CSPF suggests that fusing hidden-state representations provides a more expressive basis for preference assessment, offering a practical route toward integrated evaluative signals for non-verifiable preference tasks.

偏好评估融合方法奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。