arXiv:2603.11346cs.CVcs.GR2026-03中稿 · CVPR

用多智能体强化学习实现人与人协作动作的物理真实模仿。

Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning

论文配图:Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning
图 1 · 摘自论文原文
  • 将互助动作模仿建模为双智能体强化学习问题,联合训练辅助者与被助者策略。
  • 在基准测试中首次成功追踪复杂互助动作,实现物理真实的力交互。
  • 适合研究具身智能、人机协作或机器人服务应用的开发者参考。

类人机器人在日常服务与照护场景中具有巨大潜力。尽管近期基于物理引擎的通用运动追踪(GMT)技术已使虚拟角色和类人机器人能够复现多种人类动作,但这些行为主要局限于无接触社交互动或孤立动作。而协助场景则要求持续感知人类伙伴,并快速适应其姿态与动态变化。本文将密切互动、力交换的人类-人类动作序列模仿建模为多智能体强化学习问题。我们在物理模拟器中联合训练支持者(助手)与接受者两个智能体的伙伴感知策略,以追踪辅助动作参考。为使问题可解,我们提出一种伙伴策略初始化方案,从单人运动追踪控制器迁移先验知识,显著提升探索效率。此外,引入动态参考重定向与促进接触的奖励机制,使助手参考动作实时适配接受者的姿态,并鼓励产生物理上合理的支撑行为。实验表明,AssistMimic是首个在主流基准上成功追踪辅助交互动作的方法,验证了多智能体强化学习在物理真实且社会感知的类人控制中的优势。

原文摘要 · Abstract (English)

Humanoid robotics has strong potential to transform daily service and caregiving applications. Although recent advances in general motion tracking within physics engines (GMT) have enabled virtual characters and humanoid robots to reproduce a broad range of human motions, these behaviors are primarily limited to contact-less social interactions or isolated movements. Assistive scenarios, by contrast, require continuous awareness of a human partner and rapid adaptation to their evolving posture and dynamics. In this paper, we formulate the imitation of closely interacting, force-exchanging human-human motion sequences as a multi-agent reinforcement learning problem. We jointly train partner-aware policies for both the supporter (assistant) agent and the recipient agent in a physics simulator to track assistive motion references. To make this problem tractable, we introduce a partner policies initialization scheme that transfers priors from single-human motion-tracking controllers, greatly improving exploration. We further propose dynamic reference retargeting and contact-promoting reward, which adapt the assistant's reference motion to the recipient's real-time pose and encourage physically meaningful support. We show that AssistMimic is the first method capable of successfully tracking assistive interaction motions on established benchmarks, demonstrating the benefits of a multi-agent RL formulation for physically grounded and socially aware humanoid control.

人机协作强化学习物理模拟动作模仿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。