让机器人更懂人类协作,用稳定算法提升复杂场景下的适应能力。
HALO: Learning Human-Robot Collaboration via Heterogeneous-Agent Lyapunov Policy Optimization
- 引入李雅普诺夫约束机制,确保多智能体学习时策略更新稳定收敛。
- 在真实人形机器人实验中显著提升协作鲁棒性与泛化能力。
- 适合研究人机协同、多智能体强化学习的学者和工程师参考。
为提升人机协作(HRC)中的泛化与鲁棒性,机器人需应对多样化的交互行为与环境。现有方法面临人与机器人间固有的理性差异(理性的差距,RG),导致去中心化策略更新偏离联合优化目标。该问题本质上是可微分的非零和博弈,独立策略梯度更新可能震荡或发散。本文提出异构智能体李雅普诺夫策略优化(HALO),通过在策略参数空间施加李雅普诺夫收缩约束,稳定去中心化多智能体强化学习。不同于传统安全强化学习对状态/轨迹的约束,HALO利用李雅普诺夫认证来保障策略学习的稳定性。通过最优二次投影修正去中心化梯度,确保理性差距单调收缩,支持对开放交互空间的有效探索。大量仿真及真实人形机器人实验表明,该认证稳定性显著提升了协作性能,尤其在复杂边界情形下表现优异。项目主页:https://HaoZhang-THU.github.io/HALO/
原文摘要 · Abstract (English)
To improve generalization and resilience in human-robot collaboration (HRC), robots must contend with diverse combinations of human behaviors and contexts, motivating multi-agent reinforcement learning (MARL). However, inherent heterogeneity between robots and humans creates a rationality gap (RG), where decentralized policy updates deviate from cooperative joint optimization. The resulting learning problem is a general-sum differentiable game, so independent policy-gradient updates can oscillate or diverge without added structure. We propose heterogeneous-agent Lyapunov policy optimization (HALO), a framework that stabilizes decentralized MARL by enforcing Lyapunov-based contraction in policy-parameter space. Unlike Lyapunov-based safe RL, which targets state/trajectory constraints in constrained Markov decision processes, HALO uses Lyapunov certification to stabilize decentralized policy learning. HALO rectifies decentralized gradients via optimal quadratic projections, ensuring monotonic contraction of RG and enabling effective exploration of open-ended interaction spaces. Extensive simulations and real-world humanoid-robot experiments show that this certified stability improves generalization and robustness in collaborative corner cases. Our project website is available at https://HaoZhang-THU.github.io/HALO/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。