arXiv:2603.14741cs.CV2026-03中稿 · CVPR

让用户用点框提示精准控制被遮挡人体的补全结果。

PHAC: Promptable Human Amodal Completion

  • 用点框提示+ControlNet模块注入用户指令,实现多类型约束控制。
  • 在多个基准上生成更真实、更符合提示的补全结果,质量显著提升。
  • 适合需要精确姿态或区域控制的人体图像生成场景。

条件图像生成在以人为中心的应用中日益重要,但现有人体非完整补全(HAC)模型难以有效响应用户指定的约束,如目标姿态或空间范围。尽管姿态引导的人体图像合成(PGPIS)方法可实现显式姿态控制,却常无法保留可见部分的特定外观,且易受训练数据分布偏见影响,即使基于强大的扩散模型先验也是如此。为此,本文提出可提示人体非完整补全(PHAC),在保持可见外观的同时,根据用户提供的简单点提示(如额外关节)或框提示(如目标区域)完成遮挡人体图像。通过为每类提示设计专用ControlNet模块,将提示信号注入预训练扩散模型,并仅微调交叉注意力块以确保强提示对齐而不破坏生成先验。为进一步保留可见内容,提出基于修复的精修模块:从轻微噪声的粗略补全出发,忠实保留可见区域,并在遮挡边界实现无缝融合。在HAC和PGPIS基准上的大量实验表明,该方法生成的补全更符合物理合理性、质量更高,且相比现有方法显著提升了提示对齐能力。

原文摘要 · Abstract (English)

Conditional image generation methods are increasingly used in human-centric applications, yet existing human amodal completion (HAC) models offer users limited control over the completed content. Given an occluded person image, they hallucinate invisible regions while preserving visible ones, but cannot reliably incorporate user-specified constraints such as a desired pose or spatial extent. As a result, users often resort to repeatedly sampling the model until they obtain a satisfactory output. Pose-guided person image synthesis (PGPIS) methods allow explicit pose conditioning, but frequently fail to preserve the instance-specific visible appearance and tend to be biased toward the training distribution, even when built on strong diffusion model priors. To address these limitations, we introduce promptable human amodal completion (PHAC), a new task that completes occluded human images while satisfying both visible appearance constraints and multiple user prompts. Users provide simple point-based prompts, such as additional joints for the target pose or bounding boxes for desired regions; these prompts are encoded using ControlNet modules specialized for each prompt type. These modules inject the prompt signals into a pre-trained diffusion model, and we fine-tune only the cross-attention blocks to obtain strong prompt alignment without degrading the underlying generative prior. To further preserve visible content, we propose an inpainting-based refinement module that starts from a slightly noised coarse completion, faithfully preserves the visible regions, and ensures seamless blending at occlusion boundaries. Extensive experiments on the HAC and PGPIS benchmarks show that our approach yields more physically plausible and higher-quality completions, while significantly improving prompt alignment compared with existing amodal completion and pose-guided synthesis methods.

人体补全扩散模型提示控制图像修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。