arXiv:2601.20585cs.LGcs.AI2026-01中稿 · ICASSP2026被引 1

提出新强化学习框架,同时优化排序与回归任务。

Ranking-aware Reinforcement Learning for Ordinal Ranking

  • 统一目标联合优化回归与排序,相互促进
  • 使用可验证奖励函数评估精度与排序效果
  • 引入噪声扰动提升探索能力,避免训练停滞

序数回归与排序任务因固有的序数依赖关系而具有挑战性,传统方法难以有效建模。本文提出排名感知强化学习(RARL),一种新颖的强化学习框架,能显式学习这些依赖关系。其核心是统一目标,协同整合回归与学习排序(L2R)任务,实现两者互促提升。该目标由一个排名感知的可验证奖励驱动,联合评估回归精度与排序准确性,支持通过策略优化直接更新模型。为增强训练效果,我们引入响应变异操作(RMO),通过注入可控噪声提高探索能力,防止模型在鞍点处停滞。RARL在三个不同基准数据集上的大量实验验证了其有效性。

原文摘要 · Abstract (English)

Ordinal regression and ranking are challenging due to inherent ordinal dependencies that conventional methods struggle to model. We propose Ranking-Aware Reinforcement Learning (RARL), a novel RL framework that explicitly learns these relationships. At its core, RARL features a unified objective that synergistically integrates regression and Learning-to-Rank (L2R), enabling mutual improvement between the two tasks. This is driven by a ranking-aware verifiable reward that jointly assesses regression precision and ranking accuracy, facilitating direct model updates via policy optimization. To further enhance training, we introduce Response Mutation Operations (RMO), which inject controlled noise to improve exploration and prevent stagnation at saddle points. The effectiveness of RARL is validated through extensive experiments on three distinct benchmarks.

强化学习排序学习序数回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。