arXiv:2411.16199cs.CV2024-11CVPR被引 1

用草图和文字控制视频对象重绘,保持时序一致且精准对齐。

VIRES: Video Instance Repainting via Sketch and Text Guided Generation

  • 结合草图与文本引导,通过序列ControlNet提取结构布局。
  • 引入草图注意力机制,提升细节捕捉与语义注入能力。
  • 适用于视频编辑场景,适合需要精准控制的创作者。

我们提出VIRES,一种基于草图与文本引导的视频实例重绘方法,支持视频对象的重绘、替换、生成与移除。现有方法在时序一致性与草图序列对齐方面存在困难。VIRES利用文生视频模型的生成先验,确保时序连贯性并生成视觉美观的结果。我们提出标准化自缩放的序列ControlNet,有效提取结构布局并自适应捕捉高对比度草图细节。进一步在扩散变换器主干中引入草图注意力,以解析并注入细粒度草图语义。草图感知编码器确保重绘结果与提供的草图序列对齐。此外,我们构建了VireSet数据集,包含针对视频实例编辑任务的详细标注。实验表明,VIRES在视觉质量、时序一致性、条件对齐及人工评分上均优于当前最优方法。

原文摘要 · Abstract (English)

We introduce VIRES, a video instance repainting method with sketch and text guidance, enabling video instance repainting, replacement, generation, and removal. Existing approaches struggle with temporal consistency and accurate alignment with the provided sketch sequence. VIRES leverages the generative priors of text-to-video models to maintain temporal consistency and produce visually pleasing results. We propose the Sequential ControlNet with the standardized self-scaling, which effectively extracts structure layouts and adaptively captures high-contrast sketch details. We further augment the diffusion transformer backbone with the sketch attention to interpret and inject fine-grained sketch semantics. A sketch-aware encoder ensures that repainted results are aligned with the provided sketch sequence. Additionally, we contribute the VireSet, a dataset with detailed annotations tailored for training and evaluating video instance editing methods. Experimental results demonstrate the effectiveness of VIRES, which outperforms state-of-the-art methods in visual quality, temporal consistency, condition alignment, and human ratings. Project page: https://hjzheng.net/projects/VIRES/

视频编辑草图引导扩散模型时序一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。