arXiv:2412.12087cs.CV2024-12CVPR被引 13

用视频帧生成指令,让模型学会复杂图像操作。

Instruction-based Image Manipulation by Watching How Things Move

  • 从视频中采帧并用大模型生成编辑指令
  • 在姿态调整、元素重排等任务上表现领先
  • 适合需要自然动态的图像编辑场景

本文提出一种新颖的数据集构建流程,从视频中采样帧对,并利用多模态大语言模型(MLLMs)生成编辑指令,用于训练基于指令的图像操作模型。视频帧天然保留了主体与场景的身份信息,确保编辑过程中内容一致性。同时,视频数据捕捉了非刚性运动、复杂相机移动等多样自然动态,难以通过合成数据模拟,是构建可扩展数据集的理想来源。基于此方法,我们构建了新数据集以训练InstructMove模型,该模型能完成姿态调整、元素重排、视角变换等复杂指令式操作,在多项任务中达到当前最优性能。

原文摘要 · Abstract (English)

This paper introduces a novel dataset construction pipeline that samples pairs of frames from videos and uses multimodal large language models (MLLMs) to generate editing instructions for training instruction-based image manipulation models. Video frames inherently preserve the identity of subjects and scenes, ensuring consistent content preservation during editing. Additionally, video data captures diverse, natural dynamics-such as non-rigid subject motion and complex camera movements-that are difficult to model otherwise, making it an ideal source for scalable dataset construction. Using this approach, we create a new dataset to train InstructMove, a model capable of instruction-based complex manipulations that are difficult to achieve with synthetically generated datasets. Our model demonstrates state-of-the-art performance in tasks such as adjusting subject poses, rearranging elements, and altering camera perspectives.

图像编辑视频驱动指令学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。