解决视频编辑中生成模型失效问题,保持文本条件与空间特征
ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

- 引入目标分布正则化防止过拟合,保留生成先验
- 对比式CFG引导推理避开源视频偏差,提升编辑效果
- 异步噪声调度保护未编辑区域,适合精准视频修改场景
在预训练扩散模型上进行测试时微调(TTT)已成为强大的视频编辑范式。然而,生成模型的分布映射特性与标准TTT的单点优化之间存在根本性不匹配,导致模型出现‘先验崩溃’——丢弃文本条件和空间隐变量,使生成结果退化为源视频或不同区域特征混淆。为此,本文提出新型框架ElasticTTT,通过目标分布正则化防止尖锐记忆极小值,对比式类条件生成(Contrastive CFG)引导推理远离源视频偏见,并采用异步噪声调度保护未编辑区域。理论分析与大量实验表明,ElasticTTT成功保留了基模型的生成先验,在单次视频编辑任务上达到当前最优性能。
原文摘要 · Abstract (English)
Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-mapping nature of generative models and the single-point optimization of standard TTT. In this paper, we demonstrate that this mismatch triggers \textit{Prior Collapse}, a degenerate state where the model discards the text conditions and spatial latents, collapsing generations to the source video, or entangling the features of distinct regions. To resolve this, we propose \textbf{ElasticTTT}, a novel framework that preserves the prior generative distribution and rescues generative elasticity. Specifically, we propose \textit{Target Distribution Regularization} to prevent sharp memorization minima, \textit{Contrastive CFG} to guide inference away from source biases, and \textit{Asynchronous Noise Schedule} to preserve unedited regions. Extensive evaluations, supported by theoretical analysis, demonstrate that ElasticTTT successfully preserves the generative prior of the base model, achieving state-of-the-art performance on one-shot video editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。