让机器人像人一样灵巧操作物体,突破传统学习方法的探索瓶颈。
Towards Human-level Dexterity via Robot Learning
- 用结构化探索替代随机探索,提升强化学习效率
- 结合采样规划实现高效直接探索,显著提升灵巧操作能力
- 通过视觉触觉人类示范,实现更自然的技能模仿
灵巧智能——即使用多指手完成复杂交互的能力——是人类身体智能与高级认知能力的巅峰体现。然而,尽管表面上看似简单,人类的灵巧性实则经过数百万年的脑手协同演化才形成,包含丰富的触觉感知。实现机器人达到人类水平的灵巧性一直是机器人学的核心目标,也是迈向通用具身智能的关键里程碑。尽管计算传感运动学习已取得进展,例如实现任意物体在手中旋转,但更高阶灵巧操作仍受限于现有方法的根本缺陷。本文通过关键研究,逐步构建了一套针对多指灵巧操作的强化学习有效框架。核心在于采用结构化探索,克服强化学习中随机探索的低效问题。最终提出一种融合采样规划的直接探索强化学习方法。此外,论文还探索了基于视觉-触觉人类示范的新范式,并开发了相应的模仿学习技术。
原文摘要 · Abstract (English)
Dexterous intelligence -- the ability to perform complex interactions with multi-fingered hands -- is a pinnacle of human physical intelligence and emergent higher-order cognitive skills. However, contrary to Moravec's paradox, dexterous intelligence in humans appears simple only superficially. Many million years were spent co-evolving the human brain and hands including rich tactile sensing. Achieving human-level dexterity with robotic hands has long been a fundamental goal in robotics and represents a critical milestone toward general embodied intelligence. In this pursuit, computational sensorimotor learning has made significant progress, enabling feats such as arbitrary in-hand object reorientation. However, we observe that achieving higher levels of dexterity requires overcoming very fundamental limitations of computational sensorimotor learning. I develop robot learning methods for highly dexterous multi-fingered manipulation by directly addressing these limitations at their root cause. Chiefly, through key studies, this disseration progressively builds an effective framework for reinforcement learning of dexterous multi-fingered manipulation skills. These methods adopt structured exploration, effectively overcoming the limitations of random exploration in reinforcement learning. The insights gained culminate in a highly effective reinforcement learning that incorporates sampling-based planning for direct exploration. Additionally, this thesis explores a new paradigm of using visuo-tactile human demonstrations for dexterity, introducing corresponding imitation learning techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。