arXiv:2502.05641cs.ROcs.AI2025-02ECCV被引 12

用多模态输入生成逼真可调控的人类动作,支持不完整指令并自动补全。

Generating Physically Realistic and Directable Human Motions from Multi-Modal Inputs

论文配图:Generating Physically Realistic and Directable Human Motions from Multi-Modal Inputs
图 1 · 摘自论文原文
  • 通过掩码增强的多目标模仿学习训练人形控制器
  • 能补全缺失动作、融合多段动作、响应延迟指令
  • 适用于虚拟现实、规划系统等场景,无需微调

本工作聚焦于从多模态输入生成真实物理驱动的人类行为,这些输入可能仅部分指定所需动作。例如,输入可来自VR控制器提供的手臂运动与身体速度、部分关键点动画、视频中的计算机视觉结果,或高层级动作目标。这需要一个能处理稀疏、不完整引导的通用底层人形控制器,具备技能无缝切换和故障恢复能力。现有方法虽捕捉部分特性,但未全部实现。为此,我们提出掩码人形控制器(MHC),在增强并选择性掩码的动作演示上应用多目标模仿学习。训练后,MHC展现出追上不同步输入命令、融合多段动作序列、从稀疏多模态输入补全未指定动作的能力。我们在包含87种多样技能的数据集上训练了该模型,并展示了多种多模态应用场景,包括与规划框架集成,证明其可在无微调情况下解决用户定义的新任务。

原文摘要 · Abstract (English)

This work focuses on generating realistic, physically-based human behaviors from multi-modal inputs, which may only partially specify the desired motion. For example, the input may come from a VR controller providing arm motion and body velocity, partial key-point animation, computer vision applied to videos, or even higher-level motion goals. This requires a versatile low-level humanoid controller that can handle such sparse, under-specified guidance, seamlessly switch between skills, and recover from failures. Current approaches for learning humanoid controllers from demonstration data capture some of these characteristics, but none achieve them all. To this end, we introduce the Masked Humanoid Controller (MHC), a novel approach that applies multi-objective imitation learning on augmented and selectively masked motion demonstrations. The training methodology results in an MHC that exhibits the key capabilities of catch-up to out-of-sync input commands, combining elements from multiple motion sequences, and completing unspecified parts of motions from sparse multimodal input. We demonstrate these key capabilities for an MHC learned over a dataset of 87 diverse skills and showcase different multi-modal use cases, including integration with planning frameworks to highlight MHC's ability to solve new user-defined tasks without any finetuning.

动作生成多模态人形控制强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。