arXiv:2512.11609cs.RO2025-12被引 3

UniBYD让机器人超越模仿人类动作,自适应不同手型完成操作。

UniBYD: A Unified Framework for Learning Robotic Manipulation Across Embodiments Beyond Imitation of Human Demonstrations

  • 用统一形态表示和动态强化学习,让机器人根据自身结构自主优化抓取策略。
  • 在跨形态任务中成功率提升44.08%,远超现有方法。
  • 适合研究多形态机器人操控、需摆脱人类示范限制的场景。

在具身智能中,机器人与人类手部的形态差异导致从人类示范中学习面临巨大挑战。尽管已有研究尝试用强化学习弥补这一差距,但多数仅限于复现人类操作,性能有限,且难以支持多样化机器人手型。本文提出UniBYD——一种统一框架,通过动态强化学习算法发现与机器人物理特性匹配的操作策略。为实现对多种机器人手型的一致建模,UniBYD引入统一形态表示(UMR)。基于UMR,设计带有退火奖励机制的动态PPO,使强化学习能从离线仿效人类示范,过渡到在线自适应探索更适配多样机器人形态的策略,从而超越单纯模仿人类手部动作。针对早期策略引发的严重状态漂移问题,提出基于马尔可夫的混合影子引擎,提供细粒度引导,将模仿锚定在专家动作流形内。为评估UniBYD,构建首个跨形态操作基准UniManip,覆盖多种机器人手型。实验表明,其平均成功率较当前最先进方法提升44.08%。论文接受后将开源代码与基准数据集。

原文摘要 · Abstract (English)

In embodied intelligence, the embodiment gap between robotic and human hands brings significant challenges for learning from human demonstrations. Although some studies have attempted to bridge this gap using reinforcement learning, they remain confined to merely reproducing human manipulation, resulting in limited task performance. Moreover, current methods struggle to support diverse robotic hand configurations. In this paper, we propose UniBYD, a unified framework that uses a dynamic reinforcement learning algorithm to discover manipulation policies aligned with the robot's physical characteristics. To enable consistent modeling across diverse robotic hand morphologies, UniBYD incorporates a unified morphological representation (UMR). Building on UMR, we design a dynamic PPO with an annealed reward schedule, enabling reinforcement learning to transition from offline-informed imitation of human demonstrations to online-adaptive exploration of policies better adapted to diverse robotic morphologies, thereby going beyond mere imitation of human hands. To address the severe state drift caused by the incapacity of early-stage policies, we design a hybrid Markov-based shadow engine that provides fine-grained guidance to anchor the imitation within the expert's manifold. To evaluate UniBYD, we propose UniManip, the first benchmark for cross-embodiment manipulation spanning diverse robotic morphologies. Experiments demonstrate a 44.08% average improvement in success rate over the current state-of-the-art. Upon acceptance, we will release our code and benchmark.

机器人操控强化学习跨形态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。