arXiv:2605.14664cs.CV2026-05被引 1

用多尺度视觉语言特征实现精准视频编辑,保留原动作细节。

MiVE: Multiscale Vision-language features for reference-guided video Editing

论文配图:MiVE: Multiscale Vision-language features for reference-guided video Editing
图 1 · 摘自论文原文
  • 从多层视觉语言模型提取细粒度特征,融合到扩散Transformer中
  • 在人类偏好测试中排名第一,超越学术与商用系统
  • 适合需要精确控制视频编辑的创作者和工业级应用

参考引导的视频编辑以源视频、文本指令和参考图像为输入,要求模型忠实执行编辑指令的同时保留原始动作和未编辑内容。现有方法分为两类:解耦编码器因模态分离产生信息鸿沟;统一视觉语言编码器则因依赖最终层表征而丢失空间细节。我们发现,视觉语言模型各层编码互补信息——浅层捕捉局部空间细节,深层蕴含全局语义。基于此,提出MiVE(多尺度视觉语言特征用于参考引导视频编辑)框架,将Qwen3-VL作为多尺度特征提取器,将分层特征集成至统一自注意力扩散变换器,消除交叉注意力设计中的模态不匹配问题。实验表明,MiVE在人类偏好评估中表现最优,性能超越现有学术方法与商业系统。

原文摘要 · Abstract (English)

Reference-guided video editing takes a source video, a text instruction, and a reference image as inputs, requiring the model to faithfully apply the instructed edits while preserving original motion and unedited content. Existing methods fall into two paradigms, each with inherent limitations: decoupled encoders suffer from modality gaps when processing instructions and visual content independently, while unified vision-language encoders lose fine-grained spatial details by relying solely on final-layer representations. We observe that VLM layers encode complementary information hierarchically -- early layers capture localized spatial details essential for precise editing, while deeper layers encode global semantics for instruction comprehension. Building on this insight, we present MiVE (Multiscale Vision-language features for reference-guided video Editing), a framework that repurposes VLMs as multiscale feature extractors. MiVE extracts hierarchical features from Qwen3-VL and integrates them into a unified self-attention Diffusion Transformer, eliminating the modality mismatch inherent in cross-attention designs. Experiments demonstrate that MiVE achieves state-of-the-art performance by ranking highest in human preference, outperforming both academic methods and commercial systems.

视频编辑多模态扩散模型视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。