arXiv:2606.11525cs.ROcs.LG2026-06

通过感知物体交互动态,提升机器人抓取与操控的样本效率和成功率。

Learning Object Manipulation from Scratch via Contrastive Interaction

论文配图:Learning Object Manipulation from Scratch via Contrastive Interaction
图 1 · 摘自论文原文
  • 引入交互加权重采样机制,聚焦交互前后状态变化。
  • 在仿真中平均提升19.8%性能,真实世界击球成功率从25%升至60%。
  • 适合需要高精度交互控制的机器人任务,如抓取、推拉等。

对比强化学习(CRL)在多种目标导向的机器人任务中表现优异,但在交互密集的操控任务中仍面临挑战。本文认为,物体中心的交互(如接触或抓握)导致动态模式突变,是关键难点。我们提出将操控动态建模为分段光滑马尔可夫过程,发现交互引发的模式切换造成非线性可达性结构,标准CRL能量函数难以刻画。为此,提出交互加权重采样(IWR),在交互前、中、后阶段进行感知交互的重采样,使学习表示能保留决定未来可达性的模式边界。在2D动态控制、机器人操控及机器人冰球环境中的实验表明,IWR显著提升样本效率与整体性能,仿真平均提升19.8%。结合模拟到现实的迁移流程,使用IWR训练的策略首次实现真实世界目标导向机器人冰球击打,成功率达60%,相较之前提升35个百分点。

原文摘要 · Abstract (English)

Contrastive Reinforcement Learning (CRL) has seen recent success in a wide variety of goal-conditioned robotics tasks by learning structured representations of the dynamics. However, despite its success in locomotion and simpler control domains, CRL often struggles in interaction-rich manipulation. We argue that a key source of this difficulty is object-centric interaction, such as contact or grasping, that induces distinct changes in the underlying dynamic modes. In this work, we formulate manipulation dynamics as a piecewise-smooth Markov process and show that interaction-induced mode changes create piecewise nonlinear reachability structures that are difficult for standard CRL energy functions to represent and plan over. Based on this analysis, we introduce Interaction-weighted Resampling (IWR). IWR performs interaction-aware resampling around phases before, during, and after interactions, encouraging the learned representation to preserve the mode boundaries that determine future reachability to capture multi-modal and piecewise nonlinear reachability. Across interaction-centric environments, including 2D dynamic control, robotic manipulation, and robot air hockey, IWR improves both sample efficiency and overall performance over prior CRL methods, with 19.8% average improvement in simulation. Finally, using a sim-to-real pipeline with policies trained by IWR, we demonstrate the first real-world goal-conditioned robot air hockey agent capable of hitting goals, improving success from 25% to 60%. Project Page: IWR-arxiv.github.io.

机器人操控强化学习交互感知模拟到现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。