arXiv:2603.09882cs.ROcs.AI2026-03中稿 · Robotics: Science …

让机器人在杂乱环境中通过感知接触动态自主操作,无需人工设计规则。

Emerging Extrinsic Dexterity in Cluttered Scenes via Dynamics-aware Policy Learning

  • 基于显式世界建模学习接触引起的物体动态,用于指导强化学习。
  • 仿真中成功率比之前方法高25%以上,真实场景达50%成功率。
  • 无需手工设计接触规则,适合复杂现实场景的机器人操作任务。

外在灵巧性利用环境接触来克服抓取操作的局限。然而,在杂乱场景中实现这种灵巧性仍具挑战且研究不足,因为它需要在多个相互作用物体之间选择性地利用接触,而这些物体的动力学特性本质上是耦合的。现有方法缺乏对这种复杂动力学的显式建模,因此在杂乱环境中的非抓取操作表现不佳,限制了其在真实世界中的应用。本文提出一种动态感知策略学习(DAPL)框架,通过显式世界建模学习杂乱环境中接触诱发的物体动态表征,并将其用于条件化强化学习,从而在无需手工接触启发式或复杂奖励设计的情况下实现外在灵巧性的涌现。我们在仿真和真实世界中评估该方法。在不同密度的未见仿真杂乱场景中,该方法的成功率超过先前抓取操作、人类遥操作及基于表示的策略25%以上。真实世界中在10个杂乱场景上的成功率达到约50%,一次实际超市部署进一步验证了其良好的模拟到现实迁移能力和实用性。

原文摘要 · Abstract (English)

Extrinsic dexterity leverages environmental contact to overcome the limitations of prehensile manipulation. However, achieving such dexterity in cluttered scenes remains challenging and underexplored, as it requires selectively exploiting contact among multiple interacting objects with inherently coupled dynamics. Existing approaches lack explicit modeling of such complex dynamics and therefore fall short in non-prehensile manipulation in cluttered environments, which in turn limits their practical applicability in real-world environments. In this paper, we introduce a Dynamics-Aware Policy Learning (DAPL) framework that can facilitate policy learning with a learned representation of contact-induced object dynamics in cluttered environments. This representation is learned through explicit world modeling and used to condition reinforcement learning, enabling extrinsic dexterity to emerge without hand-crafted contact heuristics or complex reward shaping. We evaluate our approach in both simulation and the real world. Our method outperforms prehensile manipulation, human teleoperation, and prior representation-based policies by over 25% in success rate on unseen simulated cluttered scenes with varying densities. The real-world success rate reaches around 50% across 10 cluttered scenes, while a practical grocery deployment further demonstrates robust sim-to-real transfer and applicability.

机器人操作强化学习动态建模杂乱环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。