无需训练即可用图文编辑视频,保持画面连贯真实
Good Noise Makes Good Edits: A Training-Free Diffusion-Based Video Editing with Image and Text Prompts
- 通过结构化噪声图实现零样本视频编辑
- 结合参考图与文本,生成准确且连贯的视频内容
- 适合需要快速、无训练编辑视频的创作者使用
我们提出VINO,首个基于图像和文本提示的零样本、无需训练的视频编辑方法。该方法引入ρ-起始采样与稀疏双掩码机制,构建结构化噪声图,实现连贯且精准的编辑效果。为提升视觉质量,提出零图像引导策略,一种可调控的负提示方法。大量实验表明,VINO能忠实融合参考图像内容,在无需任何测试时或实例特异性训练的前提下,性能优于现有最先进方法。
原文摘要 · Abstract (English)
We propose VINO, the first zero-shot, training-free video editing method conditioned on both image and text. Our approach introduces $ρ$-start sampling and dilated dual masking to construct structured noise maps that enable coherent and accurate edits. To further enhance visual fidelity, we present zero image guidance, a controllable negative prompt strategy. Extensive experiments demonstrate that VINO faithfully incorporates the reference image into video edits, achieving strong performance compared to state-of-the-art baselines, all without any test-time or instance-specific training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。