用真实演示中的接触序列指导仿真训练,提升机器人抓取的现实迁移稳定性。
ConCent: Contact-Centric Real-to-Sim-to-Real Learning from One Demonstration

- 从真实演示中提取接触事件序列作为学习目标,引导仿真训练。
- 在接触密集任务上,迁移成功率显著高于无约束强化学习基线。
- 无需人工设计奖励函数,适合需要高物理真实性交互的任务。
从仿真到现实的策略迁移是避免大规模真实数据采集、实现机器人操作规模化的重要范式。然而,由于接触动力学差异,尤其在接触密集的任务中,微小差异可能导致任务失败。任务成功依赖于准确复现与任务相关的关键接触事件(何时、何地、如何发生)及局部接触动力学(力与运动在接触点的演化)。为此,我们提出一种接触中心的实-仿-实强化学习框架,利用真实演示中自动提取的接触事件序列作为学习目标。通过将物体近似为几何基元组,在仿真中优化其接触几何,使生成的局部接触动力学能解释观察到的状态转移。该接触序列作为结构化奖励信号,引导策略走向现实中验证过的物理合理接触模式,防止利用不真实的仿真接触。信号自动获取,无需每任务重设奖励。在多个接触密集操作任务上的实验表明,相比无约束强化学习基线,该方法实现了更稳定、鲁棒的仿真到现实迁移。
原文摘要 · Abstract (English)
Sim-to-real policy transfer -- deploying policies trained in simulation in the real world -- is a promising paradigm for scaling robot manipulation without large-scale real-world data. However, transferring simulation-trained policies remains challenging due to discrepancies in contact dynamics -- particularly in contact-rich tasks where subtle differences can alter task outcomes entirely. Because interaction between the manipulated object and the environment is mediated through contact, task success depends on accurately reproducing task-relevant contacts. Accordingly, in manipulation, contact-centric fidelity -- reproducing both the contact event sequence (when, where, and how contacts occur) and the local contact dynamics (how forces and motions evolve at each contact) -- is a necessary condition for task success. Based on this insight, we propose a contact-centric real-to-sim-to-real RL framework that uses task-relevant contact event sequences extracted from real demonstrations as the learning objective. We approximate objects as groups of primitives and optimize their contact geometry in simulation so that the resulting local contact dynamics explain the observed state transitions. The contact event sequence is automatically extracted by replaying the demonstration. This sequence serves as a structured reward signal, guiding the policy toward physically plausible contact regimes validated in reality and preventing exploitation of unrealistic simulator contacts. The signal is obtained automatically, requiring no per-task reward design. Experiments on contact-rich manipulation tasks demonstrate more stable and robust sim-to-real policy transfer compared to unconstrained RL baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。