arXiv:2412.10219cs.CV2024-12

用文本和姿态控制,把人自然地放进新场景,还能保持身份一致。

Learning Complex Non-Rigid Image Edits from Multimodal Conditioning

  • 基于稳定扩散,结合文本与姿态实现精细图像编辑
  • 在真实场景中保持人物身份,提升人与物体交互的自然度
  • 自动生成带动作差异描述的文本,降低数据标注成本

本文聚焦于将给定的人像(仅一张人物图像)插入到新场景中。我们的方法在稳定扩散基础上构建,生成自然且可控的图像,支持文本和姿态双重调控。为此,需训练成对图像:第一张为包含目标人物的参考图,第二张为同一人物不同姿势或背景的目标图。还需提供一段文本描述新姿态与参考姿态的差异。本文提出一个符合此标准的新数据集,通过人体中心、动作丰富的视频帧对,并利用多模态大模型自动生成姿态差异的文本描述。实验表明,在真实复杂场景中保持身份一致性更具挑战性,尤其当人物与物体存在交互时。结合噪声文本的弱监督与鲁棒的2D姿态信息,显著提升了人物-物体交互的质量。

原文摘要 · Abstract (English)

In this paper we focus on inserting a given human (specifically, a single image of a person) into a novel scene. Our method, which builds on top of Stable Diffusion, yields natural looking images while being highly controllable with text and pose. To accomplish this we need to train on pairs of images, the first a reference image with the person, the second a "target image" showing the same person (with a different pose and possibly in a different background). Additionally we require a text caption describing the new pose relative to that in the reference image. In this paper we present a novel dataset following this criteria, which we create using pairs of frames from human-centric and action-rich videos and employing a multimodal LLM to automatically summarize the difference in human pose for the text captions. We demonstrate that identity preservation is a more challenging task in scenes "in-the-wild", and especially scenes where there is an interaction between persons and objects. Combining the weak supervision from noisy captions, with robust 2D pose improves the quality of person-object interactions.

图像编辑姿态控制多模态身份保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。