arXiv:2601.08097cs.CLcs.LG2026-01ACL被引 1

动态调整评分策略,让模型更精准判断人类偏好。

AdaJudge: Adaptive Multi-Perspective Judging for Reward Modeling

  • 用门控模块优化模型表示,使其更适合区分细微偏好。
  • 自适应多视角池化使评分过程能灵活整合不同证据。
  • 在多个评测集上优于现有主流评分模型,适合对齐人类偏好任务。

奖励建模对对齐大语言模型与人类偏好至关重要,但现有架构普遍采用静态池化策略将序列压缩为标量分数。该范式存在两大缺陷:静态归纳偏置与任务依赖的偏好信号不匹配,以及表征错配——主干模型优化于生成任务,其表征难以支持细粒度判别。为此,我们提出AdaJudge,一个统一框架,联合优化表示与聚合过程。AdaJudge首先通过门控精炼模块将主干表示转化为判别导向空间;再以自适应多视角池化模块替代静态读出层,动态路由并融合证据。在RM-Bench和JudgeBench上的大量实验表明,AdaJudge显著优于主流奖励模型及传统池化基线。

原文摘要 · Abstract (English)

Reward modeling is essential for aligning large language models with human preferences, yet predominant architectures rely on a static pooling strategy to condense sequences into scalar scores. This paradigm, however, suffers from two key limitations: a static inductive bias that misaligns with task-dependent preference signals, and a representational mismatch, as the backbone's optimization for generation leaves its representations ill-suited to fine-grained discrimination. To address this, we propose AdaJudge, a unified framework that jointly adapts representation and aggregation. AdaJudge first improves backbone representations into a discrimination-oriented space via gated refinement blocks. It then replaces the static readout with an adaptive multi-view pooling module, which dynamically routes and combines evidence. Extensive experiments on RM-Bench and JudgeBench show that AdaJudge outperforms strong off-the-shelf reward models and traditional pooling baselines.

奖励建模人类对齐自适应池化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。