通过拆分文本与条件机制,让文字生成视频更生动有动作。
Enhancing Motion in Text-to-Video Generation with Decomposed Encoding and Conditioning
- 将文本编码拆分为静态内容和动态运动两部分处理
- 在多个数据集上显著提升视频动作自然度,保持画质清晰
- 适合需要精细动作控制的视频生成研究者使用
尽管文本到视频(T2V)生成技术不断进步,但生成具有真实运动的视频仍具挑战。现有模型常产生静态或动作微弱的输出,难以捕捉文本描述的复杂运动。问题根源在于文本编码存在内在偏差,忽略运动信息,且T2V模型的条件机制不足。为此,我们提出一种名为DEMO的新框架,通过将文本编码与条件机制分别分解为内容与运动成分,来增强运动合成。方法包含用于静态元素的内容编码器、用于时间动态的运动编码器,以及独立的内容与运动条件机制。关键创新在于引入文本-运动与视频-运动监督,提升模型对运动的理解与生成能力。在MSR-VTT、UCF-101、WebVid-10M、EvalCrafter和VBench等基准测试中,DEMO展现出更强的运动动态生成能力,同时保持高质量视觉表现。该方法通过直接从文本描述中整合全面的运动理解,显著推动了T2V生成的发展。
原文摘要 · Abstract (English)
Despite advancements in Text-to-Video (T2V) generation, producing videos with realistic motion remains challenging. Current models often yield static or minimally dynamic outputs, failing to capture complex motions described by text. This issue stems from the internal biases in text encoding, which overlooks motions, and inadequate conditioning mechanisms in T2V generation models. To address this, we propose a novel framework called DEcomposed MOtion (DEMO), which enhances motion synthesis in T2V generation by decomposing both text encoding and conditioning into content and motion components. Our method includes a content encoder for static elements and a motion encoder for temporal dynamics, alongside separate content and motion conditioning mechanisms. Crucially, we introduce text-motion and video-motion supervision to improve the model's understanding and generation of motion. Evaluations on benchmarks such as MSR-VTT, UCF-101, WebVid-10M, EvalCrafter, and VBench demonstrate DEMO's superior ability to produce videos with enhanced motion dynamics while maintaining high visual quality. Our approach significantly advances T2V generation by integrating comprehensive motion understanding directly from textual descriptions. Project page: https://PR-Ryan.github.io/DEMO-project/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。