arXiv:2607.01743cs.CV2026-07中稿 · ECCV

用分块因果扩散模型生成双人交互动作,更真实且稳定。

InterCMDM: Block-Causal Diffusion for Autoregressive Human Interaction Generation

论文配图:InterCMDM: Block-Causal Diffusion for Autoregressive Human Interaction Generation
图 1 · 摘自论文原文
  • 分两路建模每人动作因果,统一注意力捕捉互动关系。
  • 在InterHuman和Inter-X数据集上优于现有方法,长序列更连贯。
  • 通过切换注意力掩码可控制同步、反应、主从等不同互动模式。

文本条件下的双人交互生成需同时捕捉个体内部的长时序因果关系与双方紧密耦合的协调行为。现有交互扩散模型通常使用双向注意力去噪完整序列,掩盖了因果性,阻碍流式与长时生成。自回归方法虽保持因果性,但常出现时间漂移,导致协调性退化与动态不稳定。本文提出InterCMDM,一种用于自回归双人交互生成的分块因果潜变量扩散框架。InterCMDM引入双流因果扩散变换器,为每个人维护独立因果流,通过统一双流注意力与多任务掩码建模人际依赖。这些掩码在单一注意力机制中统一实现多种协调行为建模,包括同步动作、反应式响应、主从关系与独立运动。通过在不同掩码配置下联合训练作为数据增强,模型可在推理时仅通过选择对应掩码即可实现可控生成。此外,分块扩散目标支持长序列稳定潜变量演化,无需重复解码-编码循环。InterCMDM在InterHuman与Inter-X数据集上达到当前最优性能,显著提升文本-动作对齐度、真实感与长时连续性。

原文摘要 · Abstract (English)

Text-conditioned human interaction generation must capture both long-range temporal causality within each individual and tightly coupled coordination between partners. Existing interaction diffusion models typically denoise full sequences using bidirectional attention, which obscures causality and hinders streaming and long-horizon generation. Autoregressive alternatives enforce causality but often suffer from temporal drift, leading to coordination degradation and unstable interaction dynamics over time. We propose InterCMDM, a block-causal latent diffusion framework for autoregressive two-person interaction generation. InterCMDM introduces a Dual-Stream Causal Diffusion Transformer that maintains separate causal streams for each person while modeling inter-person dependencies via unified dual-stream attention with multi-task attention masks. These masks unify interaction modeling within a single attention mechanism and support diverse coordination behaviors, including simultaneous actions, reactive responses, leader-follower dynamics, and independent motion. By training a single model across these mask configurations as a form of data augmentation, InterCMDM enables controllable interaction generation by simply selecting the desired attention mask at inference time. Finally, a block-wise diffusion objective enables stable latent rollout over long sequences without repeated decode-encode cycles. InterCMDM achieves state-of-the-art performance on InterHuman and Inter-X, improving text-motion alignment, realism, and long-horizon continuity.

动作生成扩散模型因果建模双人交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。