arXiv:2505.04052cs.GRcs.CV2025-05

让插入人物自然融入场景,能自动处理遮挡并精确控制姿势。

Person-In-Situ: Scene-Consistent Human Image Insertion with Occlusion-Aware Pose Control

  • 用3D人体模型控制姿势,通过潜在扩散模型生成合适深度的人像。
  • 两种方法均在场景一致性、遮挡处理和姿势准确度上超越现有方法。
  • 适合需要精准人物合成的影视广告与虚拟拍摄场景。

将人物图像合成到场景图像中在娱乐和广告等领域有广泛应用。然而,现有方法难以处理插入人物被前景物体遮挡的问题,且常将其置于最前层,导致不自然。同时,对人物姿态的控制能力有限。为此,我们提出两种方法:均通过3D身体模型实现显式姿态控制,并利用潜在扩散模型在上下文合适的深度生成人物,自然处理遮挡,无需遮挡掩码。第一种为两阶段方法:先通过监督学习预测包含人物的场景深度图,再据此合成人物;第二种方法隐式学习遮挡,直接从输入数据生成人物,无需显式深度标注。定量与定性评估表明,两种方法在保持场景一致性、准确反映遮挡和用户指定姿态方面均优于现有方法。

原文摘要 · Abstract (English)

Compositing human figures into scene images has broad applications in areas such as entertainment and advertising. However, existing methods often cannot handle occlusion of the inserted person by foreground objects and unnaturally place the person in the frontmost layer. Moreover, they offer limited control over the inserted person's pose. To address these challenges, we propose two methods. Both allow explicit pose control via a 3D body model and leverage latent diffusion models to synthesize the person at a contextually appropriate depth, naturally handling occlusions without requiring occlusion masks. The first is a two-stage approach: the model first learns a depth map of the scene with the person through supervised learning, and then synthesizes the person accordingly. The second method learns occlusion implicitly and synthesizes the person directly from input data without explicit depth supervision. Quantitative and qualitative evaluations show that both methods outperform existing approaches by better preserving scene consistency while accurately reflecting occlusions and user-specified poses.

图像合成姿态控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。