通过噪声扰动概念嵌入,提升视频生成的动态效果。
MotionCFG: Boosting Motion Dynamics via Stochastic Concept Perturbation
- 用高斯噪声扰动概念嵌入,生成隐式负样本。
- 在多个SOTA模型上提升运动动态性,计算开销极小。
- 适合需要精细动作控制或复杂概念引导的场景。
尽管文本到视频(T2V)合成取得进展,但生成高保真且动态的运动仍具挑战。现有方法主要依赖无分类器指导(CFG)和显式负向提示(如"静态"、"模糊")抑制不良伪影,但常引入语义偏差并破坏物体完整性,称为内容-运动漂移。为此,我们提出MotionCFG,通过对比目标概念与其噪声扰动版本,增强运动动态性。具体地,将高斯噪声注入概念嵌入,生成局部负向锚点,涵盖广泛的次优运动变化空间。相比显式否定,该方法实现隐式硬负样本挖掘,不改变全局语义身份,聚焦于时序细节优化。结合分段指导调度,仅在去噪早期干预,可在多个SOTA T2V框架中持续提升运动动态性,计算开销可忽略,视觉质量损失极小。此外,该噪声诱导的对比机制不仅适用于锐化运动轨迹,还能有效调控精确物体数量等复杂非线性概念,而这些通常难以通过标准文本引导调节。
原文摘要 · Abstract (English)
Despite recent advances in Text-to-Video (T2V) synthesis, generating high-fidelity and dynamic motion remains a significant challenge. Existing methods primarily rely on Classifier-Free Guidance (CFG), often with explicit negative prompts (e.g. "static", "blurry"), to suppress undesired artifacts. However, such explicit negations frequently introduce unintended semantic bias and distort object integrity; a phenomenon we define as Content-Motion Drift. To address this, we propose MotionCFG, a framework that enhances motion dynamics by contrasting a target concept with its noise-perturbed counterparts. Specifically, by injecting Gaussian noise into the concept embeddings, MotionCFG creates localized negative anchors that encapsulate a broad complementary space of sub-optimal motion variations. Unlike explicit negations, this approach facilitates implicit hard negative mining without shifting the global semantic identity, allowing for a focused refinement of temporal details. Combined with a piecewise guidance schedule that confines intervention to the early denoising steps, MotionCFG consistently improves motion dynamics across state-of-the-art T2V frameworks with negligible computational overhead and minimal compromise in visual quality. Additionally, we demonstrate that this noise-induced contrastive mechanism is effective not only for sharpening motion trajectories but also for steering complex, non-linear concepts such as precise object numerosity, which are typically difficult to modulate via standard text-based guidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。