arXiv:2602.08557cs.RO2026-02被引 2

通过约束采样与强化学习结合,提升机器人复杂接触环境下的非抓取操作能力。

Combined Constrained Sampling and Reinforcement Learning for Robotic Manipulation

  • 设计基于接触结构的约束状态采样器,增强探索多样性。
  • 在接触丰富的任务中实现90%以上成功率,优于传统重置方法。
  • 适合需要通用、动态非抓取操作的机器人研究者使用。

在高接触场景下训练非抓取操作策略是机器人学的核心挑战。尽管强化学习(RL)在此类场景中表现出色,但仍难以充分探索并发现复杂的操作策略。为此,我们结合两种基本思路:首先,设计合理的重置策略(即每轮起点分布)已被证明能提升RL的探索效率;其次,虽然基于模型的方法在寻找操作轨迹上困难,但近期研究表明,在约束流形上采样状态可极为高效。基于此,我们提出一种新型状态采样器,显著提升了目标条件强化学习在复杂接触丰富任务中的性能。该采样器显式考虑了接触结构,以覆盖多样化的接触模式。通过结合约束采样重置、投影插值与课程学习,新方法在无需手动设计动作序列的情况下,成功训练出通用、非抓取且动态的操纵策略,在接触密集环境中达到90%以上的成功率,显著优于无约束采样及其它重置策略。更多信息见 https://www.user.tu-berlin.de/mtoussai/26-CSRL/

原文摘要 · Abstract (English)

Training non-prehensile manipulation policies in contact-rich settings is a core challenge in robotics. While Reinforcement Learning (RL) has demonstrated its strength in such settings, it may struggle to sufficiently explore and discover complex manipulation strategies. To address this, we combine two basic ideas: First, designing appropriate reset strategies (the start state distribution of episodes) has shown promise in improving RL exploration and effectiveness. Second, while model-based approaches to finding trajectories through manipulation are hard, recent work showed that model-based approaches to sampling states on constrained manifolds can be highly efficient. Based on these observations, we propose a novel state sampler that boosts the performance of goal-conditioned RL in complex contact-rich manipulation tasks. Our sampler explicitly takes into account the structure of contact in order to provide a rich covering of diverse contact modes. By combining constrained sampling resets with projected interpolation and curriculum learning, our novel approach outperforms RL without constrained sampling and alternative reset methods, and effectively trains universal, non-prehensile, and dynamic manipulation policies in contact-rich settings. See https://www.user.tu-berlin.de/mtoussai/26-CSRL/ for supplementary material.

强化学习机器人操控接触建模状态采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。