arXiv:2607.02034cs.CV2026-07

让虚拟人物在复杂场景中自然互动,同时保持动作真实。

ComplexMimic: Human-Scene Interaction Imitation in Complex 3D Environments

论文配图:ComplexMimic: Human-Scene Interaction Imitation in Complex 3D Environments
图 1 · 摘自论文原文
  • 双专家策略:一个模仿动作,一个处理碰撞适应
  • 难样本优先训练,提升复杂行为学习效果
  • 适用于需要真实物理交互的虚拟角色生成

基于物理的人-场景交互(HSI)模仿学习对具身智能至关重要,它连接了三维动作与真实世界动态。然而,现有方法多聚焦于简化场景,限制了其在真实环境中的应用。本文聚焦复杂环境下的HSI模仿。我们发现,在复杂环境中存在动作完成度与动作自然性之间的固有权衡。为此,提出ComplexMimic框架,通过解析不完美动捕数据重建多样化的交互行为。首先引入双流策略,学习两个互补专家:一个用于精确运动追踪,另一个用于复杂场景中的碰撞感知适应。其次,传统多专家蒸馏平权分配监督信号,常导致挑战性行为欠采样,影响学习效率。为此,提出一种难度感知蒸馏策略,基于失败统计和学习进度信号,自适应加权并优先学习难但可学的轨迹。在三个基准数据集上的大量实验表明,本方法优于当前最优技术。

原文摘要 · Abstract (English)

Physics-based Human-Scene Interaction (HSI) imitation learning is crucial for embodied intelligence as it bridges the gap between kinematic 3D motions and real-world dynamics. However, most existing methods focus on simplified scene settings, leaving complex environments largely unexplored, which limits their applicability in real-world scenarios. In this paper, we focus on HSI mimicry in complex environments. Under this complex setting, we observe an inherent trade-off between successfully performing interaction and maintaining natural, physically plausible motions. To address this challenge, we propose ComplexMimic, a framework that reconstructs diverse HSI by interpreting imperfect MoCap data. First, we introduce a Dual Flow Strategy, which learns two complementary experts: an imitation expert for accurate motion tracking and an interaction expert for collision-aware adaptation in complex scenes. Second, naive multi-expert distillation, which treats all experts equally, often under-samples challenging behaviors, limiting effective learning. To mitigate this issue, we propose a difficulty-aware distillation strategy that adaptively weights supervision and prioritizes hard-yet-learnable trajectories guided by failure statistics and learning progress signals. Extensive experiments on three benchmark datasets demonstrate that our approach outperforms current state-of-the-art methods.

人机交互动作模仿物理模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。