无需训练,通过关键层激活信息实现文本控制的视频编辑。
TV-LiVE: Training-Free, Text-Guided Video Editing via Layer Informed Vitality Exploitation
- 利用模型中关键层的激活信息,精准定位可编辑区域。
- 在添加新物体和非刚性变形任务上优于现有方法。
- 适合需要快速、可控视频修改的研究者与创作者。
随着基于扩散模型的视频生成技术快速发展,视频编辑受到越来越多关注。当前研究多集中于风格迁移、背景替换、物体替换和属性修改等任务,保持源视频内容结构不变。然而,新增物体或非刚性形变等复杂任务仍缺乏探索。本文提出无需训练的文本引导视频编辑框架TV-LiVE,通过层感知活力挖掘实现高效编辑。我们实证发现,视频生成模型中特定层对输出质量影响显著,且与旋转位置编码(RoPE)密切相关。基于此,方法通过有选择地将源模型的关键和值特征注入目标模型对应层,实现对象添加与非刚性编辑。对于对象添加,进一步识别显著层以提取对应新提示的目标掩码区域,实验表明这些掩码能准确指示需编辑区域。结果证明,TV-LiVE在两类任务上均优于现有方法。
原文摘要 · Abstract (English)
Video editing has garnered increasing attention alongside the rapid progress of diffusion-based video generation models. As part of these advancements, there is a growing demand for more accessible and controllable forms of video editing, such as prompt-based editing. Previous studies have primarily focused on tasks such as style transfer, background replacement, object substitution, and attribute modification, while maintaining the content structure of the source video. However, more complex tasks, including the addition of novel objects and nonrigid transformations, remain relatively unexplored. In this paper, we present TV-LiVE, a Training-free and text-guided Video editing framework via Layerinformed Vitality Exploitation. We empirically identify vital layers within the video generation model that significantly influence the quality of generated outputs. Notably, these layers are closely associated with Rotary Position Embeddings (RoPE). Based on this observation, our method enables both object addition and non-rigid video editing by selectively injecting key and value features from the source model into the corresponding layers of the target model guided by the layer vitality. For object addition, we further identify prominent layers to extract the mask regions corresponding to the newly added target prompt. We found that the extracted masks from the prominent layers faithfully indicate the region to be edited. Experimental results demonstrate that TV-LiVE outperforms existing approaches for both object addition and non-rigid video editing. Project Page: https://emjay73.github.io/TV_LiVE/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。