arXiv:2503.06897cs.CV2025-03

用去冗余融合提升文本驱动动作生成的精度与细节。

Multi-granular body modeling with Redundancy-Free Spatiotemporal Fusion for Text-Driven Motion Generation

  • 并行处理局部与整体身体建模,捕捉精细关节运动。
  • 双向时间模块捕获短期细节与长期依赖,提升动作连贯性。
  • 动态融合模块消除冗余,增强时空表征,适合动画与游戏应用。

文本到动作生成结合多模态学习与计算机图形学,可简化游戏、动画、机器人和虚拟现实的内容创作。现有方法通常简单堆叠空间与时间特征,引入冗余且忽略细微关节线索。本文提出HiSTF Mamba框架,包含双空间Mamba、双时间Mamba和动态时空融合模块(DSFM)。双空间模块并行运行局部与整体身体模型,同时捕捉整体协调与精细关节运动;双时间模块双向扫描序列,编码短期细节与长期依赖;DSFM去除冗余时间信息,提取互补线索并与空间特征融合,构建更丰富的时空表示。在HumanML3D基准测试中,该方法在多个指标上表现优异,实现高保真度与文本-动作间的紧密语义对齐。

原文摘要 · Abstract (English)

Text-to-motion generation sits at the intersection of multimodal learning and computer graphics and is gaining momentum because it can simplify content creation for games, animation, robotics and virtual reality. Most current methods stack spatial and temporal features in a straightforward way, which adds redundancy and still misses subtle joint-level cues. We introduce HiSTF Mamba, a framework with three parts: Dual-Spatial Mamba, Bi-Temporal Mamba and a Dynamic Spatiotemporal Fusion Module (DSFM). The Dual-Spatial module runs part-based and whole-body models in parallel, capturing both overall coordination and fine-grained joint motion. The Bi-Temporal module scans sequences forward and backward to encode short-term details and long-term dependencies. DSFM removes redundant temporal information, extracts complementary cues and fuses them with spatial features to build a richer spatiotemporal representation. Experiments on the HumanML3D benchmark show that HiSTF Mamba performs well across several metrics, achieving high fidelity and tight semantic alignment between text and motion.

动作生成文本驱动时空融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。