用多模态大模型统一指导图像视频编辑,效果领先。
InstructX: Towards Unified Visual Editing with MLLM Guidance
- 用图像数据训练的多模态模型可自发掌握视频编辑能力
- 融合模态特异性特征,单模型实现图像视频统一编辑
- 无需视频标注也能在复杂任务上表现优异
随着多模态大语言模型(MLLM)在视觉理解与推理方面取得进展,利用其提升扩散模型的编辑性能成为研究热点。然而,现有工作缺乏对MLLM设计选择的深入分析,且在复杂任务如视频编辑中,二者融合仍面临挑战。本文提出InstructX,一个统一的图像与视频编辑框架。我们系统研究了基于指令驱动的MLLM与扩散模型集成策略,涵盖多样化任务。实验表明:(1) 在图像数据上训练的模型能涌现视频编辑能力,无需显式视频监督,缓解视频数据稀缺问题;(2) 通过引入模态特异性MLLM特征,实现图像与视频编辑任务的统一建模。大量实验验证该方法可处理广泛编辑任务,并达到当前最优性能。
原文摘要 · Abstract (English)
With recent advances in Multimodal Large Language Models (MLLMs) showing strong visual understanding and reasoning, interest is growing in using them to improve the editing performance of diffusion models. Despite rapid progress, most studies lack an in-depth analysis of MLLM design choices. Moreover, the integration of MLLMs and diffusion models remains an open challenge in some difficult tasks, such as video editing. In this paper, we present InstructX, a unified framework for image and video editing. Specifically, we conduct a comprehensive study on integrating MLLMs and diffusion models for instruction-driven editing across diverse tasks. Building on this study, we analyze the cooperation and distinction between images and videos in unified modeling. (1) We show that training on image data can lead to emergent video editing capabilities without explicit supervision, thereby alleviating the constraints imposed by scarce video training data. (2) By incorporating modality-specific MLLM features, our approach effectively unifies image and video editing tasks within a single model. Extensive experiments demonstrate that our method can handle a broad range of image and video editing tasks and achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。