arXiv:2502.20390cs.CVcs.GR2025-02CVPR被引 89

用一套策略学习多种人物交互,从粗糙数据中生成逼真动作。

InterMimic: Towards Universal Whole-Body Control for Physics-Based Human-Object Interactions

  • 分阶段训练:先学特定人,再提炼成通用策略。
  • 在多个数据集上实现零样本泛化,动作自然多样。
  • 适合做虚拟角色动画或物理仿真,尤其对数据不完美场景友好。

实现人类与多种物体的逼真物理交互模拟是长期目标。将基于物理的动作模仿扩展至复杂人-物交互(HOI)面临巨大挑战,包括人与物体间复杂的耦合关系、物体几何形状的多样性以及动作捕捉数据中的缺陷,如接触信息不准、手部细节不足等。我们提出 InterMimic 框架,使单一策略能够从覆盖多种动态、多样的全身体交互的数小时不完美动捕数据中稳健学习。核心思路是‘先完美,再扩展’:首先为每个个体训练专用教师策略,用于模仿、重定向和精炼动捕数据;随后将这些教师策略的知识蒸馏到学生策略中,教师作为在线专家提供直接监督和高质量参考。值得注意的是,我们对学生策略进行强化学习微调,使其超越简单示范复制,获得更高质量的解。实验表明,InterMimic 在多个 HOI 数据集上均能生成真实且多样的交互动作,所学策略可零样本泛化,并无缝集成于运动学生成器,使框架从单纯模仿跃升为复杂人-物交互的生成建模。

原文摘要 · Abstract (English)

Achieving realistic simulations of humans interacting with a wide range of objects has long been a fundamental goal. Extending physics-based motion imitation to complex human-object interactions (HOIs) is challenging due to intricate human-object coupling, variability in object geometries, and artifacts in motion capture data, such as inaccurate contacts and limited hand detail. We introduce InterMimic, a framework that enables a single policy to robustly learn from hours of imperfect MoCap data covering diverse full-body interactions with dynamic and varied objects. Our key insight is to employ a curriculum strategy -- perfect first, then scale up. We first train subject-specific teacher policies to mimic, retarget, and refine motion capture data. Next, we distill these teachers into a student policy, with the teachers acting as online experts providing direct supervision, as well as high-quality references. Notably, we incorporate RL fine-tuning on the student policy to surpass mere demonstration replication and achieve higher-quality solutions. Our experiments demonstrate that InterMimic produces realistic and diverse interactions across multiple HOI datasets. The learned policy generalizes in a zero-shot manner and seamlessly integrates with kinematic generators, elevating the framework from mere imitation to generative modeling of complex human-object interactions.

人-物交互动作模仿生成建模强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。