arXiv:2607.09866cs.ROcs.AI2026-07被引 1

提升机器人离线到在线强化学习中的价值估计可靠性,实现更稳定高效的策略优化。

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning

论文配图:Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning
图 1 · 摘自论文原文
  • 构建历史条件化的价值估计器,用全局进展与局部偏好评估可靠性。
  • 可靠价值函数使离线策略预训练效果优于无质量感知的行为克隆,提升在线优化稳定性。
  • 适用于需要高精度和泛化能力的机器人操作任务,如芯片插入与积木拆解。

离线到在线强化学习在通用机器人操作中具有前景,但其全栈复杂性阻碍了复现与诊断。价值估计在优先处理异构数据以改进策略中起核心作用。然而,价值函数可靠性如何影响策略优化这一核心问题仍缺乏研究。为此,我们提出 Robo-ValueRL,一个统一框架,实现可靠的价值估计,并系统追踪其对策略预训练与在线改进的影响。具体地,Robo-ValueRL 学习历史条件化的价值估计器,并通过全局进展与局部偏好指标评估其可靠性。这些价值估计被用于质量条件化的共识策略预训练和在线回放中的残差适应模块,形成统一测试平台以分析价值可靠性对下游性能的影响。在240小时离线演示和超过3,000条在线回放轨迹上,实验表明下游性能与价值可靠性密切相关:可靠价值函数提供更优的动作质量估计,使价值引导的离线强化学习比无质量感知的行为克隆更具可扩展性,并通过优先高质量回放数据稳定在线改进。结合可靠价值引导的离线预训练与在线优化,系统在毫米级精密芯片插入任务上达到86%成功率,在通用积木拆解任务上达84%成功率。研究强调了价值引导的数据利用对有效策略改进的重要性。

原文摘要 · Abstract (English)

Offline-to-online reinforcement learning is promising for generalizable robotic manipulation, yet its full-stack complexity obscures reproduction and diagnosis. Within such systems, value estimation plays a central role in prioritizing heterogeneous data for policy improvement. Despite its importance, the central question remains underexplored: how value-function reliability shapes policy optimization in offline-to-online reinforcement learning. To answer this question, we propose Robo-ValueRL, a unified framework that enables reliable value estimation and systematically traces its downstream effects on policy pretraining and online improvement. Concretely, Robo-ValueRL learns a history-conditioned value estimator and evaluates its reliability through global-progress and local-preference metrics. These resulting value estimates are propagated into quality-conditioned consistency-policy pretraining and a residual adaptation module on online rollouts, providing a unified testbed for analyzing how value reliability shapes downstream policy performance. Across 240 hours of offline demonstrations and over 3,000 online rollout trajectories, our extensive experiments show that downstream performance is strongly associated with value reliability. Reliable value functions provide better action-quality estimates, allowing value-guided offline RL to scale more effectively than quality-agnostic behavior cloning, and stabilize online improvement by prioritizing high-quality rollout data. Integrating reliable value guidance through offline pretraining with online improvement, our system achieves 86% success on millimeter-level precise chip insertion and 84% on generalizable block disassembly. We hope these findings highlight the importance of value-guided data utilization for effective policy improvement from heterogeneous robotic experience.

强化学习机器人操作价值估计离线到在线

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。