arXiv:2510.00144cs.LGcs.AI2025-10被引 1

在奖励有限时,选对样本能大幅降低标注成本

Which Rewards Matter? Reward Selection for Reinforcement Learning under Limited Feedback

  • 根据状态访问和价值函数选择最有影响力的样本标注
  • 用少量标注实现接近全量标注的策略性能
  • 适合资源受限下的强化学习应用

强化学习的效果取决于训练期间可用的奖励。但在实际问题中,由于计算或财务限制,尤其是依赖人工反馈时,获取大量奖励标签往往不可行。当只能对部分样本进行奖励标注时,关键问题是:应选择哪些样本标注以最大化策略性能?本文提出了奖励选择问题的正式框架(RLLF),研究了两类策略:(i) 基于无奖励信息(如状态访问频率、部分价值函数)的启发式方法;(ii) 使用辅助评估反馈预训练的策略。研究发现,有效的奖励集中在两类样本上:(1) 引导智能体走向最优轨迹的样本;(2) 在偏离后帮助恢复近最优行为的样本。有效选择方法在远少于全监督所需标注数的情况下,仍能获得接近最优的策略,证明奖励选择是反馈受限场景下强化学习可扩展的关键范式。

原文摘要 · Abstract (English)

The ability of reinforcement learning algorithms to learn effective policies is determined by the rewards available during training. However, for practical problems, obtaining large quantities of reward labels is often infeasible due to computational or financial constraints, particularly when relying on human feedback. When reinforcement learning must proceed with limited feedback -- only a fraction of samples get rewards labeled -- a fundamental question arises: which samples should be labeled to maximize policy performance? We formalize this problem of reward selection for reinforcement learning from limited feedback (RLLF), introducing a new problem formulation that facilitates the study of strategies for selecting impactful rewards. Two types of selection strategies are investigated: (i) heuristics that rely on reward-free information such as state visitation and partial value functions, and (ii) strategies pre-trained using auxiliary evaluative feedback. We find that critical subsets of rewards are those that (1) guide the agent along optimal trajectories, and (2) support recovery toward near-optimal behavior after deviations. Effective selection methods yield near-optimal policies with significantly fewer reward labels than full supervision, establishing reward selection as a powerful paradigm for scaling reinforcement learning in feedback-limited settings.

强化学习奖励选择少样本反馈受限

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。