arXiv:2511.19435cs.CV2025-11被引 4

用视频模型零样本编辑图片,让静态图像动起来。

Are Image-to-Video Models Good Zero-Shot Image Editors?

  • 把指令变成带时间线索的推理提示,让模型理解动态变化
  • 删减冗余帧特征,提速同时保持画面连贯性
  • 最后一步锐化模糊帧,适合需要精准修改的用户

大规模视频扩散模型展现出强大的世界模拟和时序推理能力,但其作为零样本图像编辑器的应用仍不充分。本文提出IF-Edit,一种无需微调的框架,将预训练的图像到视频扩散模型用于指令驱动的图像编辑。该方法解决三个核心挑战:提示错位、冗余时序隐变量、后期帧模糊。具体包括:(1) 链式思维提示增强模块,将静态编辑指令转化为时序对齐的推理提示;(2) 时序隐变量丢弃策略,在专家切换点后压缩帧隐变量,加速去噪并保持语义与时序一致性;(3) 自洽后处理优化步骤,通过短时静帧轨迹锐化晚期帧。在四个公开基准测试上,涵盖非刚性编辑、物理与时序推理、通用指令编辑任务,结果表明IF-Edit在以推理为核心的任务中表现优异,且在通用编辑任务中保持竞争力。本研究系统揭示了视频扩散模型作为图像编辑器的潜力,并提出统一视频-图像生成推理的简单方案。

原文摘要 · Abstract (English)

Large-scale video diffusion models show strong world simulation and temporal reasoning abilities, but their use as zero-shot image editors remains underexplored. We introduce IF-Edit, a tuning-free framework that repurposes pretrained image-to-video diffusion models for instruction-driven image editing. IF-Edit addresses three key challenges: prompt misalignment, redundant temporal latents, and blurry late-stage frames. It includes (1) a chain-of-thought prompt enhancement module that transforms static editing instructions into temporally grounded reasoning prompts; (2) a temporal latent dropout strategy that compresses frame latents after the expert-switch point, accelerating denoising while preserving semantic and temporal coherence; and (3) a self-consistent post-refinement step that sharpens late-stage frames using a short still-video trajectory. Experiments on four public benchmarks, covering non-rigid editing, physical and temporal reasoning, and general instruction edits, show that IF-Edit performs strongly on reasoning-centric tasks while remaining competitive on general-purpose edits. Our study provides a systematic view of video diffusion models as image editors and highlights a simple recipe for unified video-image generative reasoning.

图像编辑视频生成扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。