无需标签,自动识别机器人操作中的成功与失败片段。
STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning

- 用专家示范数据自监督训练时序预测器集合,生成帧级优势分数。
- 在真实机器人任务中准确识别停滞、失败和恢复行为,提升成功率59%以上。
- 适合需要从复杂演示中提取有效信息的机器人学习场景。
现实世界机器人学习依赖异构数据,但示范与回放常混杂有效进展、停滞、修正及次优行为。高效策略学习需区分可靠局部进展与失败。本文提出自监督时序集成优势建模(STEAM),从专家示范中无标签学习此类优势。STEAM 在专家轨迹的帧对上训练一组时序偏移预测器,以两帧间的归一化时序偏移作为自监督信号。每个预测器将帧对映射为时序偏移分布,并转换为标量优势值。最终取集成中最小优势值,保守评估混合质量回放数据。在真实双臂毛巾折叠、芯片结账、可乐补货及单臂抓取-放置任务中,STEAM 能准确识别停滞、失败与恢复。结合CFGRL后,成功率分别较基线提升59%、54.3%、23%和16.2%。
原文摘要 · Abstract (English)
Real-world robot learning increasingly relies on heterogeneous data, but demonstrations and rollouts often mix useful progress with stalls, corrections, and suboptimal behavior. Effective policy learning therefore requires frame-level advantages that distinguish reliable local progress from failures and regressions. We propose Self-supervised Temporal Ensemble Advantage Modeling (STEAM), a label-free method that learns such advantages from expert demonstrations. STEAM trains an ensemble of temporal-offset predictors on frame pairs within expert trajectories, using the normalized temporal offset between two frames as a self-supervised signal. Each predictor maps a frame pair to a distribution over temporal offsets, which is converted into a scalar advantage. STEAM then takes the minimum advantage across the ensemble to score mixed-quality rollout data conservatively. Across real-world bimanual towel folding, chip checkout, cola restocking, and single-arm pick-and-place tasks, STEAM identifies stalls, failures, and recoveries. When combined with CFGRL, STEAM further improves policy success rate by 59%, 54.3%, 23% and 16.2% over baselines, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。