arXiv:2606.15918cs.RO2026-06

用强化学习让机器人手臂更省电,一次充电多抓几次果子。

Energy-Efficient Arm Reaching for a Humanoid Robot via Deep Reinforcement Learning with Identified Power Models

  • 结合物理电能模型与SAC算法,用奖励函数平衡动作精度和能耗。
  • 仿真中成功率达69.9%,平均每次抓取耗能98.16焦耳。
  • 实机验证效果稳定,误差在允许范围内,适合能源受限的机器人任务。

人形机器人在实地操作任务(如采摘苹果)中面临严重能耗限制,直接影响单次充电下的可达动作次数。本文提出一种面向Unitree G1人形机器人左臂(7自由度)的端到端、能耗感知强化学习框架,将基于物理的实验识别电能模型与在Pinocchio模拟器中训练的Soft Actor-Critic(SAC)策略相结合。策略采用增量关节位置动作空间,并使用混合星座奖励函数,融合四点末端执行器构型距离与扭矩范数能耗代理。经过5×10⁶次训练,在动力学仿真中对1000个随机目标实现69.9%的成功率,成功轨迹平均能耗为98.16焦耳。在真实Unitree G1上,对三批各10个目标独立测试,平均能耗71.5±48.3焦耳,末端位置误差2.64±1.04厘米,姿态误差6.92±1.33°,均在4厘米/8.6°的训练容忍范围内。该成果为能耗感知强化学习在人形机器人手臂运动中迈出关键一步。

原文摘要 · Abstract (English)

Humanoid robots performing in-field manipulation tasks, such as robotic apple harvesting, face severe energy constraints that directly limit the number of reaching motions that can be executed per battery charge. This paper presents an end-to-end, energy-aware reinforcement learning framework for the 7-degree-of-freedom left arm of the Unitree~G1 humanoid robot, combining a physics-based, experimentally identified electrical power model with a Soft Actor-Critic (SAC) policy trained in a Pinocchio-based rigid-body dynamics simulator. The RL policy operates on an incremental joint-position action space and is trained with a Hybrid Constellation Reward that combines a four-point end-effector constellation distance with a torque-norm energy proxy; after % $5\times10^6$ training it reaches a $69.9\%$ success rate over $1\,000$ random targets in kinematic simulation, at a mean energy of \SI{98.16}{\joule} on successful episodes. Finally, on the physical Unitree~G1, the policy is validated over three independent 10-target batches, achieving a mean energy of $71.5 \pm 48.3$\,J, an end-effector position error of $2.64 \pm 1.04$\,cm, and an orientation error of $6.92 \pm 1.33^\circ$ -- within the \SI{4}{\centi\metre}/$8.6^\circ$ training tolerance. These results constitute a first step toward energy-aware reinforcement-learning-based arm reaching for humanoid robots.

强化学习机器人能耗人形机器人能量效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。