用多尺度视觉语言特征实现精准视频编辑,保留原动作细节。
MiVE: Multiscale Vision-language features for reference-guided video Editing

- 从多层视觉语言模型提取细粒度特征,融合到扩散Transformer中
- 在人类偏好测试中排名第一,超越学术与商用系统
- 适合需要精确控制视频编辑的创作者和工业级应用
参考引导的视频编辑以源视频、文本指令和参考图像为输入,要求模型忠实执行编辑指令的同时保留原始动作和未编辑内容。现有方法分为两类:解耦编码器因模态分离产生信息鸿沟;统一视觉语言编码器则因依赖最终层表征而丢失空间细节。我们发现,视觉语言模型各层编码互补信息——浅层捕捉局部空间细节,深层蕴含全局语义。基于此,提出MiVE(多尺度视觉语言特征用于参考引导视频编辑)框架,将Qwen3-VL作为多尺度特征提取器,将分层特征集成至统一自注意力扩散变换器,消除交叉注意力设计中的模态不匹配问题。实验表明,MiVE在人类偏好评估中表现最优,性能超越现有学术方法与商业系统。
原文摘要 · Abstract (English)
Reference-guided video editing takes a source video, a text instruction, and a reference image as inputs, requiring the model to faithfully apply the instructed edits while preserving original motion and unedited content. Existing methods fall into two paradigms, each with inherent limitations: decoupled encoders suffer from modality gaps when processing instructions and visual content independently, while unified vision-language encoders lose fine-grained spatial details by relying solely on final-layer representations. We observe that VLM layers encode complementary information hierarchically -- early layers capture localized spatial details essential for precise editing, while deeper layers encode global semantics for instruction comprehension. Building on this insight, we present MiVE (Multiscale Vision-language features for reference-guided video Editing), a framework that repurposes VLMs as multiscale feature extractors. MiVE extracts hierarchical features from Qwen3-VL and integrates them into a unified self-attention Diffusion Transformer, eliminating the modality mismatch inherent in cross-attention designs. Experiments demonstrate that MiVE achieves state-of-the-art performance by ranking highest in human preference, outperforming both academic methods and commercial systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。