arXiv:2605.11369cs.CV2026-05

用预训练模块组合生成长时间人物交互动态动作。

Dynamic Full-body Motion Agent with Object Interaction via Blending Pre-trained Modular Controllers

论文配图:Dynamic Full-body Motion Agent with Object Interaction via Blending Pre-trained Modular Controllers
图 1 · 摘自论文原文
  • 用扩散模型增强数据并规划动态交互序列
  • 组合不同专家模型提升成功率与交互持续性
  • 适合需要高效生成复杂交互动作的研究者

生成真实物理特性的人物交互动态动作仍具挑战,因现有数据集多限于静态交互,而预训练代理仅能处理无物体的动态全身动作或静态交互。近期工作如InsActor和CLoSD虽在规划与执行阶段生成交互动作,但仅支持静态或短时接触(如击打)。本文提出框架,通过结合预训练运动先验与模仿代理,在规划与执行阶段实现长时间、动态的交互动作(如持桌奔跑)。规划阶段利用预训练人体运动扩散模型引入动态先验,并生成物体轨迹,以规划动态交互序列;执行阶段采用编排网络融合专精于动态人体动作或静态交互动作的预训练模仿代理,实现时空上的技能互补。所提方法在多个动态交互任务中持续提升成功率,同时保持交互连贯性。消融实验验证了数据增强与编排融合的有效性。相较于基线,本方法显著减少训练时间且性能相当。

原文摘要 · Abstract (English)

Generating physically plausible dynamic motions of human-object interaction (HOI) remains challenging, mainly due to existing HOI datasets limited to static interactions, and pretrained agents capable of either dynamic full-body motions without objects or static HOI motions. Recent works such as InsActor and CLoSD generate HOI motions in planning and execution stages, are yet limited to either static or short-term contacts e.g. striking. In this work, we propose a framework that fulfills dynamic and long-term interaction motions such as running while holding a table, by combining pretrained motion priors and imitation agents in planning and execution stages. In the planning stage, we augment HOI datasets with dynamic priors from a pretrained human motion diffusion model, followed by object trajectory generation. This plans dynamic HOI sequences. In the execution stage, a composer network blends actions of pretrained imitation agents specialized either for dynamic human motions or static HOI motions, enabling spatio-temporal composition of their complementary skills. Our method over relevant prior-arts consistently improves success rates while maintaining interaction for dynamic HOI tasks. Furthermore, blending pretrained experts with our composer achieves competitive performance in significantly reduced training time. Ablation studies validate the effectiveness of our augmentation and composer blending.

动作生成人机交互扩散模型模块化控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。