arXiv:2608.16745cs.CV2026-08

用视觉例子指导视频编辑,让模型更懂复杂动作和细节。

VicEdit: Learning to Edit Videos from Visual In-Context Examples

论文配图:VicEdit: Learning to Edit Videos from Visual In-Context Examples
图 1 · 摘自论文原文
  • 用图像、视频对等视觉样本替代纯文字指令。
  • 在10类任务上达当前最佳,40万样本数据集支持。
  • 适合需要精准控制视频编辑的创作者或研究者。

尽管指令式视频编辑取得进展,但单一文本指令难以表达精细纹理和复杂动态。为此,我们提出视觉上下文编辑新范式,将视频编辑从文本指令升级为包含单图、图像对和视频对的多模态视觉引导。为支持该范式,我们构建了首个大规模视觉上下文视频编辑数据集VicEdit-400K,通过自动化流水线生成40万条高质量样本,覆盖十类任务,利用多维过滤保障视觉保真度与语义一致性。基于此,我们提出VicEdit统一框架,通过模态自适应语义蒸馏,从异构视觉参考中提取特定模态语义标记,并结合双上下文注入机制,使文本与视觉信号协同作用于生成过程。在VicEditBench上的广泛评估表明,VicEdit在基础指令编辑与视觉上下文编辑任务中均达到最先进性能,确立视觉上下文学习作为强大且可控的视频编辑范式。

原文摘要 · Abstract (English)

Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptual gap, we propose Visual In-context Editing, a new paradigm elevating video editing from textual instructions to multi-modal visual guidance encompassing single image, image pair, and video pair. To facilitate this paradigm, we curate VicEdit-400K, the first large-scale dataset for visual in-context video editing. We develop an automated pipeline to generate 400K high-quality samples across ten task types, ensuring superior visual fidelity and semantic consistency through multi-dimensional filtering. Leveraging this foundation, we introduce VicEdit, a unified framework to bridge visual and textual contexts. To adaptively extract editing semantics from heterogeneous references, we design Modality-Adaptive Semantic Distillation, which produces modality-specific semantic tokens from visual references. These tokens are then synergistically integrated with textual instructions through Dual-Context Injection, enabling the generation process to benefit from both visual and textual signals. Extensive evaluations on VicEditBench demonstrate that VicEdit achieves state-of-the-art performance across both basic instruction editing and visual in-context editing tasks, establishing visual in-context learning as a powerful and controllable paradigm for video editing.

视频编辑多模态视觉引导生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。