arXiv:2510.06915cs.CLcs.AI2025-10被引 5

让奖励模型学会长上下文一致性判断,提升AI对复杂对话的理解能力

LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling

  • 设计多阶段训练法,使模型能跨长上下文保持偏好判断一致
  • 8B模型性能超越70B大模型,接近商业级Gemini 2.5 Pro
  • 新基准测试揭示现有模型在长对话中易失效,适合对话系统研发者

奖励模型(RM)在对齐大语言模型与人类偏好中起关键作用。随着真实应用越来越多涉及长历史轨迹(如智能体交互),评估模型响应不仅需高质量,还需与上下文保持一致。然而,现有奖励模型仍局限于短上下文场景,主要关注回复层面属性(如安全或帮助性),忽视了长上下文-响应一致性这一关键维度。本文提出Long-RewardBench,一个专为长上下文奖励模型评估设计的基准,包含成对比较和Best-of-N任务。初步研究发现,即使最先进的生成式奖励模型在长上下文场景下也表现出显著脆弱性,难以维持上下文感知的偏好判断。基于对模型输出失败模式的分析,我们提出一种通用的多阶段训练策略,可有效将任意模型扩展为鲁棒的长上下文奖励模型(LongRM)。实验表明,该方法不仅显著提升长上下文评估表现,同时保留强短上下文能力。值得注意的是,我们的8B LongRM性能超越更大规模的70B基线模型,并达到商用模型Gemini 2.5 Pro水平。

原文摘要 · Abstract (English)

Reward model (RM) plays a pivotal role in aligning large language model (LLM) with human preferences. As real-world applications increasingly involve long history trajectories, e.g., LLM agent, it becomes indispensable to evaluate whether a model's responses are not only high-quality but also grounded in and consistent with the provided context. Yet, current RMs remain confined to short-context settings and primarily focus on response-level attributes (e.g., safety or helpfulness), while largely neglecting the critical dimension of long context-response consistency. In this work, we introduce Long-RewardBench, a benchmark specifically designed for long-context RM evaluation, featuring both Pairwise Comparison and Best-of-N tasks. Our preliminary study reveals that even state-of-the-art generative RMs exhibit significant fragility in long-context scenarios, failing to maintain context-aware preference judgments. Motivated by the analysis of failure patterns observed in model outputs, we propose a general multi-stage training strategy that effectively scales arbitrary models into robust Long-context RMs (LongRMs). Experiments show that our approach not only substantially improves performance on long-context evaluation but also preserves strong short-context capability. Notably, our 8B LongRM outperforms much larger 70B-scale baselines and matches the performance of the proprietary Gemini 2.5 Pro model.

奖励模型长上下文对齐技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。