用多轨迹相似状态修正时间进度标签,提升机器人学习的准确性
UR-VC: Unsupervised Robotic Value Correction for Time-Derived Progress Proxies

- 通过检索不同任务中相似状态,聚合其时间标签生成更准进度估计
- 在真实双臂布料折叠任务中,能捕捉局部退步与非均匀进展
- 无需人工标注或额外模型,适合长时序柔体操作任务研究者
现代机器人学习系统依赖密集的进展或价值信号来评估中间状态、指导策略学习并检测任务完成,这些信号的质量至关重要。由于密集标签难以大规模获取,演示中的归一化时间常被用作可扩展替代:越晚的帧视为更高进展。但这种时间衍生标签仅是物理进展的噪声代理。在接触丰富的操作任务中,机器人可能先取得进展,随后因滑动、抓取失败或部分撤销而损失成果,而时间标签仍持续单调上升。本文提出无监督机器人价值修正(UR-VC),一种离线、无需训练的方法,用于修正时间衍生的进展标签。UR-VC利用演示数据中的一个简单规律:相似状态在不同任务周期中重复出现,但对应不同时间戳。不依赖单条轨迹的时间戳,而是从其他周期检索相似状态,并聚合其时间标签以获得修正后的进展估计。该方法无需人工进展标签、奖励标注或额外价值模型。我们在真实双臂布料平整折叠数据集上进行评估,这是一个具有明显中间进展的长时序柔体操作任务。修正后的标签能够捕捉局部退步和非均匀进展,而归一化时间无法表示;同时保留整体任务趋势。进一步将修正信号用于构建优势标签,支持近期基于优势条件的策略学习框架。在相同数据、模型与训练设置下,UR-VC在真实机器人任务成功率上表现出正向趋势。
原文摘要 · Abstract (English)
Modern robot learning systems increasingly rely on dense progress or value signals to evaluate intermediate states, guide policy learning, and detect task completion, making the quality of these signals critical. Since such dense labels are rarely available at scale, normalized time within a demonstration is often used as a scalable substitute: later frames are treated as higher progress. However, this time-derived label is only a noisy proxy for physical task progress. In contact-rich manipulation, a robot may make progress and then lose it through slips, failed grasps, or partial undoing, while the time-derived label continues to increase monotonically. We introduce Unsupervised Robotic Value Correction (UR-VC), an offline, training-free method for correcting time-derived progress labels. UR-VC exploits a simple regularity in demonstration data: similar states often recur across different episodes, but at different timestamps. Instead of trusting the timestamp from a single trajectory, UR-VC retrieves similar states from other episodes and aggregates their time-derived labels to obtain a corrected progress estimate. UR-VC requires no manual progress labels, reward annotations, or additional value model. We evaluate UR-VC on real bimanual cloth flatten-and-fold data, a long-horizon deformable-object manipulation task with visible intermediate progress. The corrected labels capture local regressions and non-uniform progress that normalized time cannot represent, while preserving the overall task trend. We further use the corrected signal to construct advantage labels for VLA training, following recent advantage-conditioned policy learning. UR-VC shows a positive trend in real-robot task success under matched data, model, and training settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。