arXiv:2505.09827cs.CV2025-05CVPR被引 5

用状态空间模型生成任意长度双人交互动作,解决长序列运动合成难题。

Dyadic Mamba: Long-term Dyadic Human Motion Synthesis

  • 基于状态空间模型,通过拼接实现双人动作信息高效传递。
  • 在长序列上显著优于基于Transformer的方法,短序列表现也具竞争力。
  • 提出新基准评估长期动作合成质量,助力后续研究。

从文本描述生成逼真的双人交互动作面临挑战,尤其在超出典型训练序列长度的长期交互中。尽管近期基于Transformer的方法在短时合成上表现良好,但受限于位置编码机制,在长序列上表现不佳。本文提出Dyadic Mamba,一种利用状态空间模型(SSMs)生成任意长度高质量双人动作的新方法。其简单有效的架构通过拼接实现个体动作序列间的信息流动,无需复杂交叉注意力机制。实验表明,Dyadic Mamba在标准短时基准上表现竞争力,且在长序列上显著优于基于Transformer的方法。此外,我们提出了一个新基准以评估长期动作合成质量,为未来研究提供标准化框架。结果表明,基于SSM的架构为解决从文本生成长期双人动作这一难题提供了有前景的方向。

原文摘要 · Abstract (English)

Generating realistic dyadic human motion from text descriptions presents significant challenges, particularly for extended interactions that exceed typical training sequence lengths. While recent transformer-based approaches have shown promising results for short-term dyadic motion synthesis, they struggle with longer sequences due to inherent limitations in positional encoding schemes. In this paper, we introduce Dyadic Mamba, a novel approach that leverages State-Space Models (SSMs) to generate high-quality dyadic human motion of arbitrary length. Our method employs a simple yet effective architecture that facilitates information flow between individual motion sequences through concatenation, eliminating the need for complex cross-attention mechanisms. We demonstrate that Dyadic Mamba achieves competitive performance on standard short-term benchmarks while significantly outperforming transformer-based approaches on longer sequences. Additionally, we propose a new benchmark for evaluating long-term motion synthesis quality, providing a standardized framework for future research. Our results demonstrate that SSM-based architectures offer a promising direction for addressing the challenging task of long-term dyadic human motion synthesis from text descriptions.

动作生成状态空间模型双人交互长序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。