用视频帧生成指令,让模型学会复杂图像操作。
Instruction-based Image Manipulation by Watching How Things Move
- 从视频中采帧并用大模型生成编辑指令
- 在姿态调整、元素重排等任务上表现领先
- 适合需要自然动态的图像编辑场景
本文提出一种新颖的数据集构建流程,从视频中采样帧对,并利用多模态大语言模型(MLLMs)生成编辑指令,用于训练基于指令的图像操作模型。视频帧天然保留了主体与场景的身份信息,确保编辑过程中内容一致性。同时,视频数据捕捉了非刚性运动、复杂相机移动等多样自然动态,难以通过合成数据模拟,是构建可扩展数据集的理想来源。基于此方法,我们构建了新数据集以训练InstructMove模型,该模型能完成姿态调整、元素重排、视角变换等复杂指令式操作,在多项任务中达到当前最优性能。
原文摘要 · Abstract (English)
This paper introduces a novel dataset construction pipeline that samples pairs of frames from videos and uses multimodal large language models (MLLMs) to generate editing instructions for training instruction-based image manipulation models. Video frames inherently preserve the identity of subjects and scenes, ensuring consistent content preservation during editing. Additionally, video data captures diverse, natural dynamics-such as non-rigid subject motion and complex camera movements-that are difficult to model otherwise, making it an ideal source for scalable dataset construction. Using this approach, we create a new dataset to train InstructMove, a model capable of instruction-based complex manipulations that are difficult to achieve with synthetically generated datasets. Our model demonstrates state-of-the-art performance in tasks such as adjusting subject poses, rearranging elements, and altering camera perspectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。