通过双模态联合建模,提升复杂人体动作视频生成质量
EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
- 采用双分支结构融合视频与动作数据,统一建模外观与运动关系
- 在8万对视频-动作数据上训练,显著提升动作连贯性与合理性
- 适合需要高质量人体动作生成的研究者与开发者使用
视频生成模型虽有长足进步,但因人体关节自由度高,仍难以生成复杂动作。这源于仅基于像素的训练目标会过度关注外观保真度,而忽略运动规律。为此,我们提出EchoMotion框架,联合建模外观与人体运动分布,以提升复杂人体动作视频生成质量。该框架扩展了DiT(Diffusion Transformer)架构,采用双分支结构处理来自不同模态的拼接标记;并提出MVS-RoPE(Motion-Video Synchronized RoPE),为视频与运动标记提供统一的3D位置编码,建立时间对齐的归纳偏置。我们还设计了两阶段训练策略,支持复杂人体动作视频及对应运动序列的联合生成,以及跨模态条件生成任务。为支持训练,构建了约8万条高质量、以人为中心的视频-动作配对数据集HuMoVe。结果表明,显式表示人体运动可与外观信息互补,显著提升生成视频的连贯性与合理性。
原文摘要 · Abstract (English)
Video generation models have advanced significantly, yet they still struggle to synthesize complex human movements due to the high degrees of freedom in human articulation. This limitation stems from the intrinsic constraints of pixel-only training objectives, which inherently bias models toward appearance fidelity at the expense of learning underlying kinematic principles. To address this, we introduce EchoMotion, a framework designed to model the joint distribution of appearance and human motion, thereby improving the quality of complex human action video generation. EchoMotion extends the DiT (Diffusion Transformer) framework with a dual-branch architecture that jointly processes tokens concatenated from different modalities. Furthermore, we propose MVS-RoPE (Motion-Video Syncronized RoPE), which offers unified 3D positional encoding for both video and motion tokens. By providing a synchronized coordinate system for the dual-modal latent sequence, MVS-RoPE establishes an inductive bias that fosters temporal alignment between the two modalities. We also propose a Motion-Video Two-Stage Training Strategy. This strategy enables the model to perform both the joint generation of complex human action videos and their corresponding motion sequences, as well as versatile cross-modal conditional generation tasks. To facilitate the training of a model with these capabilities, we construct HuMoVe, a large-scale dataset of approximately 80,000 high-quality, human-centric video-motion pairs. Our findings reveal that explicitly representing human motion is complementary to appearance, significantly boosting the coherence and plausibility of human-centric video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。