arXiv:2508.15669cs.ROcs.LG2025-08中稿 · IROS 2025被引 1

通过检测并扰动机械臂停顿状态,提升灵巧操作的鲁棒性。

Exploiting Policy Idling for Dexterous Manipulation

  • 在检测到动作停滞时施加扰动,帮助策略跳出低效区域。
  • 真实世界插入任务成功率提升15%-35%,无需额外训练。
  • 适合需要高精度灵巧操作的研究者和工程师参考。

近年来,基于学习的灵巧操作方法取得了显著进展,但其策略仍缺乏可靠性,对关键变化因素的鲁棒性有限。一个常见失败模式是策略出现‘停顿’现象,即在某些状态下停止运动,仅在小范围内徘徊。这种现象往往源于训练数据中的微小动作,例如抓取或插入物体前的高精度操作阶段。已有工作尝试通过过滤数据或调整控制频率缓解,但可能损害其他性能。本文提出一种新方法——暂停诱导扰动(Pause-Induced Perturbations, PIP),在检测到停顿状态时施加扰动,引导策略逃离不利吸引域。在多种复杂双臂模拟任务中,该方法显著提升测试性能,且无需额外监督或训练。由于机器人常在动作关键点停顿,由此生成的轨迹也更有利于迭代策略优化。在真实世界多指插入任务中,成功率绝对提升15%-35%。

原文摘要 · Abstract (English)

Learning-based methods for dexterous manipulation have made notable progress in recent years. However, learned policies often still lack reliability and exhibit limited robustness to important factors of variation. One failure pattern that can be observed across many settings is that policies idle, i.e. they cease to move beyond a small region of states when they reach certain states. This policy idling is often a reflection of the training data. For instance, it can occur when the data contains small actions in areas where the robot needs to perform high-precision motions, e.g., when preparing to grasp an object or object insertion. Prior works have tried to mitigate this phenomenon e.g. by filtering the training data or modifying the control frequency. However, these approaches can negatively impact policy performance in other ways. As an alternative, we investigate how to leverage the detectability of idling behavior to inform exploration and policy improvement. Our approach, Pause-Induced Perturbations (PIP), applies perturbations at detected idling states, thus helping it to escape problematic basins of attraction. On a range of challenging simulated dual-arm tasks, we find that this simple approach can already noticeably improve test-time performance, with no additional supervision or training. Furthermore, since the robot tends to idle at critical points in a movement, we also find that learning from the resulting episodes leads to better iterative policy improvement compared to prior approaches. Our perturbation strategy also leads to a 15-35% improvement in absolute success rate on a real-world insertion task that requires complex multi-finger manipulation.

灵巧操作强化学习机器人控制策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。