arXiv:2607.14681cs.CV2026-07被引 2

让多参考图视频编辑更准:用显式绑定解决来源混淆问题

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships

论文配图:ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships
图 1 · 摘自论文原文
  • 引入带语义位置的参考标记,明确视觉属性与来源的对应关系
  • 在多个数据集上优于通用大模型,开源方法中效果领先
  • 适合需要精准控制多参考图像输入的视频生成研究者

基于扩散的视频生成模型在多参考图像条件视频编辑方面取得显著进展。然而,现有方法仍难以准确协调多个视觉源的信息。我们识别出关键缺陷:现有编辑指令缺乏显式的参考关系,大多数多模态大语言模型无法可靠生成此类关系。为此,我们提出ReBind,一个系统性框架,采用嵌入参考标记的语义指令作为多参考图像条件视频编辑的中间表示。核心洞察是将参考标记嵌入语义位置,消除歧义并建立视觉属性与来源间的精确绑定。我们开发了ReBind-Instruct,一种专用的MLLM,通过两阶段渐进式方案学习建立视觉属性与参考源之间的显式绑定。进一步设计ReBind-Edit,实现文本到视频模型的轻量级适配,使模型能将视觉属性绑定到指定来源。大量实验表明,ReBind在指令质量上显著优于通用大模型,并在参考图像条件视频编辑任务中达到开源方法的顶尖水平。

原文摘要 · Abstract (English)

Recent diffusion-based video generation models have made significant progress in multi-reference image-conditioned video editing. However, existing methods still struggle to coordinate information from multiple visual sources accurately. We identify a critical deficiency in existing approaches. Existing editing instructions lack explicit reference relationships, and most multimodal large language models (MLLMs) cannot generate them reliably. To address this problem, we propose ReBind, a systematic framework that introduces semantic instructions with embedded reference tokens as the intermediate representation for multi-reference image-conditioned video editing. Our key insight is embedding reference tokens at semantic positions to eliminate ambiguity and establish precise bindings between visual attributes and their sources. We develop ReBind-Instruct, a specialized MLLM that learns to establish explicit bindings between visual attributes and their reference sources through a two-stage progressive scheme for precise reference relationships. We further develop ReBind-Edit, which enables lightweight adaptation of text-to-video models to coordinate multiple references by binding visual attributes to their designated sources. Extensive experiments demonstrate that ReBind substantially outperforms general-purpose MLLMs in instruction quality and achieves state-of-the-art performance among open-source methods on reference image conditioned video editing. Our project webpage: https://rebind-mrv2v.github.io/.

视频编辑多参考扩散模型大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。