arXiv:2508.19575cs.CVcs.AI2025-08被引 3

让人物与物体的互动生成更精准可控,支持身份保留和姿势调节。

Interact-Custom: Customized Human Object Interaction Image Generation

  • 分拆身份特征与交互姿态特征,通过两阶段生成实现精细控制。
  • 在自建数据集上生成交互图像,保持人物与物体身份一致且姿态自然。
  • 适合需要高精度人物-物体交互生成的应用场景,如影视设计、虚拟试穿。

组合式定制图像生成旨在对生成内容中的多个目标概念进行定制,应用广泛。现有方法多关注目标实体外观的保留,却忽视了实体间的细粒度交互控制。为此,本文聚焦人-物交互场景,提出定制化人-物交互图像生成(CHOI)任务,要求同时保留目标人物与物体的身份,并精确控制二者之间的交互语义。该任务面临两大挑战:(1) 身份保留与交互控制需将人-物分解为独立的身份特征与姿态相关的交互特征,但现有HOI图像数据集缺乏适合此类特征解耦学习的样本;(2) 人物与物体间不合理的空间配置会导致交互语义缺失。为此,我们构建了一个大规模数据集,其中同一对人-物包含不同交互姿态的样本。随后设计两阶段模型Interact-Custom:首先生成描绘交互行为的前景掩码以显式建模空间配置,再在掩码引导下生成交互图像,同时保留身份特征。此外,若用户提供背景图及目标位置,模型还可额外指定这些条件,实现更高程度的内容可控性。在专为CHOI任务设计的指标上,大量实验验证了方法的有效性。

原文摘要 · Abstract (English)

Compositional Customized Image Generation aims to customize multiple target concepts within generation content, which has gained attention for its wild application. Existing approaches mainly concentrate on the target entity's appearance preservation, while neglecting the fine-grained interaction control among target entities. To enable the model of such interaction control capability, we focus on human object interaction scenario and propose the task of Customized Human Object Interaction Image Generation(CHOI), which simultaneously requires identity preservation for target human object and the interaction semantic control between them. Two primary challenges exist for CHOI:(1)simultaneous identity preservation and interaction control demands require the model to decompose the human object into self-contained identity features and pose-oriented interaction features, while the current HOI image datasets fail to provide ideal samples for such feature-decomposed learning.(2)inappropriate spatial configuration between human and object may lead to the lack of desired interaction semantics. To tackle it, we first process a large-scale dataset, where each sample encompasses the same pair of human object involving different interactive poses. Then we design a two-stage model Interact-Custom, which firstly explicitly models the spatial configuration by generating a foreground mask depicting the interaction behavior, then under the guidance of this mask, we generate the target human object interacting while preserving their identities features. Furthermore, if the background image and the union location of where the target human object should appear are provided by users, Interact-Custom also provides the optional functionality to specify them, offering high content controllability. Extensive experiments on our tailored metrics for CHOI task demonstrate the effectiveness of our approach.

图像生成人物交互可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。