arXiv:2601.20383cs.CV2026-01被引 4

提出首个自回归多人运动生成框架,实现动态人数与长文本的精准生成。

HINT: Hierarchical Interaction Modeling for Autoregressive Multi-Human Motion Generation

  • 分层交互建模:解耦局部动作与人之间互动,适配不同人数
  • 滑动窗口策略提升效率,保持长序列连贯性与细节精度
  • 支持变长输入与动态人数,适合复杂场景的实时生成应用

基于文本驱动的多人复杂交互运动生成仍具挑战。现有离线方法受限于固定长度和固定人数,难以应对长文本或可变参与人数。为此,本文提出HINT,首个基于扩散模型的自回归多人运动生成框架,引入分层交互建模机制。HINT采用解耦的运动表示,在标准化潜空间中分离局部动作语义与人之间交互,无需额外微调即可适应不同人数。通过滑动窗口策略实现高效在线生成,聚合窗口内局部与跨窗口全局条件,捕捉历史轨迹、人与人依赖关系并匹配文本引导。该设计在长序列上兼顾细粒度交互建模与整体连贯性。在公开数据集上的实验表明,HINT性能媲美强基线离线模型,并显著优于现有自回归方法。在InterHuman数据集上,FID达3.100,较此前最优结果5.154大幅提升。

原文摘要 · Abstract (English)

Text-driven multi-human motion generation with complex interactions remains a challenging problem. Despite progress in performance, existing offline methods that generate fixed-length motions with a fixed number of agents, are inherently limited in handling long or variable text, and varying agent counts. These limitations naturally encourage autoregressive formulations, which predict future motions step by step conditioned on all past trajectories and current text guidance. In this work, we introduce HINT, the first autoregressive framework for multi-human motion generation with Hierarchical INTeraction modeling in diffusion. First, HINT leverages a disentangled motion representation within a canonicalized latent space, decoupling local motion semantics from inter-person interactions. This design facilitates direct adaptation to varying numbers of human participants without requiring additional refinement. Second, HINT adopts a sliding-window strategy for efficient online generation, and aggregates local within-window and global cross-window conditions to capture past human history, inter-person dependencies, and align with text guidance. This strategy not only enables fine-grained interaction modeling within each window but also preserves long-horizon coherence across all the long sequence. Extensive experiments on public benchmarks demonstrate that HINT matches the performance of strong offline models and surpasses autoregressive baselines. Notably, on InterHuman, HINT achieves an FID of 3.100, significantly improving over the previous state-of-the-art score of 5.154.

运动生成自回归多人交互扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。