让人物与物体自然互动,保持姿态和外观一致。
HOComp: Interaction-Aware Human-Object Composition
- 用多模态大模型识别交互区域和类型,指导生成精细姿态
- 在新数据集IHOC上,合成效果显著优于现有方法
- 适合需要真实人物-物体交互的图像生成任务
现有图像引导合成方法在插入前景物体时,难以实现自然的人物-物体交互融合。本文提出HOComp,一种面向以人物为中心背景的物体合成方法,确保前景物体与背景人物之间和谐交互及外观一致性。核心设计包括:(1)基于MLLM的分区域姿态引导(MRPG),利用多模态大模型识别交互区域与类型(如手持、左持),提供粗到细的姿态约束,并结合人体关键点追踪动作变化,施加细粒度姿态控制;(2)细节一致的外观保持(DCAP),通过形状感知注意力调制、多视角外观损失与背景一致性损失,保障前景形状/纹理一致,还原背景人物真实外观。同时构建首个交互感知人-物合成数据集IHOC。实验表明,HOComp在该数据集上可生成更自然的交互效果,定性与定量均优于现有方法。
原文摘要 · Abstract (English)
While existing image-guided composition methods may help insert a foreground object onto a user-specified region of a background image, achieving natural blending inside the region with the rest of the image unchanged, we observe that these existing methods often struggle in synthesizing seamless interaction-aware compositions when the task involves human-object interactions. In this paper, we first propose HOComp, a novel approach for compositing a foreground object onto a human-centric background image, while ensuring harmonious interactions between the foreground object and the background person and their consistent appearances. Our approach includes two key designs: (1) MLLMs-driven Region-based Pose Guidance (MRPG), which utilizes MLLMs to identify the interaction region as well as the interaction type (e.g., holding and lefting) to provide coarse-to-fine constraints to the generated pose for the interaction while incorporating human pose landmarks to track action variations and enforcing fine-grained pose constraints; and (2) Detail-Consistent Appearance Preservation (DCAP), which unifies a shape-aware attention modulation mechanism, a multi-view appearance loss, and a background consistency loss to ensure consistent shapes/textures of the foreground and faithful reproduction of the background human. We then propose the first dataset, named Interaction-aware Human-Object Composition (IHOC), for the task. Experimental results on our dataset show that HOComp effectively generates harmonious human-object interactions with consistent appearances, and outperforms relevant methods qualitatively and quantitatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。