arXiv:2603.02115cs.ROcs.AI2026-03被引 46

用轨迹对比提升机器人奖励模型泛化能力,支持失败数据学习

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons

  • 通过帧级进度+轨迹间偏好对比双目标训练,兼顾局部与全局监督
  • 在百万级轨迹数据集上训练,显著提升复杂任务中的奖励泛化性
  • 适合需要从大量失败数据中学习的机器人强化学习场景

通用机器人奖励模型通常基于专家示范预测绝对任务进展,仅提供局部帧级监督。尽管在专家示范中有效,但该范式在大规模机器人数据集中表现不佳,因失败和次优轨迹众多,密集进度标签难以定义。本文提出Robometer,一种结合轨迹内进度监督与轨迹间偏好监督的可扩展奖励建模框架。其双目标训练包括:帧级进度损失,锚定专家数据上的奖励量级;以及轨迹比较偏好损失,对同一任务的轨迹施加全局排序约束,从而有效利用真实和增强的失败轨迹。为支持该方法,我们构建了包含超百万条轨迹的RBM-1M数据集,涵盖多样机器人形态与任务,包含大量次优与失败数据。在多个基准与真实世界评估中,Robometer学习到的奖励函数比先前方法更具泛化性,并在多种下游应用中提升机器人学习性能。

原文摘要 · Abstract (English)

General-purpose robot reward models are typically trained to predict absolute task progress from expert demonstrations, providing only local, frame-level supervision. While effective for expert demonstrations, this paradigm scales poorly to large-scale robotics datasets where failed and suboptimal trajectories are abundant and assigning dense progress labels is ambiguous. We introduce Robometer, a scalable reward modeling framework that combines intra-trajectory progress supervision with inter-trajectory preference supervision. Robometer is trained with a dual objective: a frame-level progress loss that anchors reward magnitude on expert data, and a trajectory-comparison preference loss that imposes global ordering constraints across trajectories of the same task, enabling effective learning from both real and augmented failed trajectories. To support this formulation at scale, we curate RBM-1M, a reward-learning dataset comprising over one million trajectories spanning diverse robot embodiments and tasks, including substantial suboptimal and failure data. Across benchmarks and real-world evaluations, Robometer learns more generalizable reward functions than prior methods and improves robot learning performance across a diverse set of downstream applications. Code, model weights, and videos at https://robometer.github.io/.

机器人奖励模型轨迹对比强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。