arXiv:2412.04558cs.CV2024-12被引 3

让图像根据动作指令动态改变物体位置或姿势

Action-based image editing guided by human instructions

  • 通过对比动作差异学习,响应动作类文本指令
  • 在视频帧数据上训练,实现从起始场景到动作终态的生成
  • 适合需要动态姿态调整的图像编辑场景

文本驱动的图像编辑通常被视为静态任务,即根据人类指令对输入图像中的元素进行插入、删除或修改。鉴于该任务的静态特性,本文旨在使其动态化,引入动作概念。通过此方法,我们希望在保持物体视觉属性不变的前提下,改变物体的位置或姿态以表现不同动作。为实现这一挑战性任务,我们提出一种对动作文本指令敏感的新模型,通过学习识别对比性动作差异来实现。模型训练基于从展示动作前后视觉场景的视频中提取的帧构建的新数据集。实验表明,使用动作类文本指令进行图像编辑取得显著提升,并具备强大推理能力,能够将输入图像作为动作起始场景,生成展示动作终态的新图像。

原文摘要 · Abstract (English)

Text-based image editing is typically approached as a static task that involves operations such as inserting, deleting, or modifying elements of an input image based on human instructions. Given the static nature of this task, in this paper, we aim to make this task dynamic by incorporating actions. By doing this, we intend to modify the positions or postures of objects in the image to depict different actions while maintaining the visual properties of the objects. To implement this challenging task, we propose a new model that is sensitive to action text instructions by learning to recognize contrastive action discrepancies. The model training is done on new datasets defined by extracting frames from videos that show the visual scenes before and after an action. We show substantial improvements in image editing using action-based text instructions and high reasoning capabilities that allow our model to use the input image as a starting scene for an action while generating a new image that shows the final scene of the action.

图像编辑动作生成文本驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。