让视频生成更精准地匹配动作语义与视觉细节。
SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
- 分离主体与动作的语义表示,避免混淆。
- 通过可高效调整的视觉适配器提升动作真实感。
- 新训练策略与基准测试验证效果领先。
基于扩散模型的视频动作定制化生成,能从少量视频样本中获取人体动作表征,并通过精确文本控制实现任意主体迁移。现有方法依赖语义对齐,但忽略动作复杂的时空视觉特征;仅优化视觉又易造成语义混乱。为此,我们提出SynMotion,联合利用语义引导与视觉自适应。在语义层面,设计双嵌入理解机制,解耦主体与动作表征,保留多样主体生成能力;在视觉层面,将参数高效的运动适配器集成至预训练视频生成模型,增强动作保真度与时间连贯性。此外,提出一种嵌入特定的交替优化训练策略,依托手动构建的主体先验视频(SPV)数据集,提升动作特异性同时保持跨主体泛化能力。最后,构建了包含多样化动作模式的新基准MotionBench。T2V与I2V设置下的实验结果表明,该方法显著优于现有基线。
原文摘要 · Abstract (English)
Diffusion-based video motion customization facilitates the acquisition of human motion representations from a few video samples, while achieving arbitrary subjects transfer through precise textual conditioning. Existing approaches often rely on semantic-level alignment, expecting the model to learn new motion concepts and combine them with other entities (e.g., ''cats'' or ''dogs'') to produce visually appealing results. However, video data involve complex spatio-temporal patterns, and focusing solely on semantics cause the model to overlook the visual complexity of motion. Conversely, tuning only the visual representation leads to semantic confusion in representing the intended action. To address these limitations, we propose SynMotion, a new motion-customized video generation model that jointly leverages semantic guidance and visual adaptation. At the semantic level, we introduce the dual-embedding semantic comprehension mechanism which disentangles subject and motion representations, allowing the model to learn customized motion features while preserving its generative capabilities for diverse subjects. At the visual level, we integrate parameter-efficient motion adapters into a pre-trained video generation model to enhance motion fidelity and temporal coherence. Furthermore, we introduce a new embedding-specific training strategy which \textbf{alternately optimizes} subject and motion embeddings, supported by the manually constructed Subject Prior Video (SPV) training dataset. This strategy promotes motion specificity while preserving generalization across diverse subjects. Lastly, we introduce MotionBench, a newly curated benchmark with diverse motion patterns. Experimental results across both T2V and I2V settings demonstrate that \method outperforms existing baselines. Project page: https://lucaria-academy.github.io/SynMotion/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。