提出多层扩散框架,解决多人重叠场景下人像精准擦除难题
MILD: Multi-Layer Diffusion Strategy for Complex and Precise Multi-IP Aware Human Erasing
- 分层去噪路径分别重建每个主体与背景,实现空间解耦
- 在复杂遮挡场景中显著降低边界伪影,结构恢复质量提升32%
- 适合需要高精度人像处理的视频编辑、图像修复领域
近年来扩散模型在图像定制任务中表现卓越,但现有掩码引导的人像擦除方法在人-人遮挡、人-物纠缠及人-背景干扰等复杂场景下仍存在不足,主要源于缺乏大规模多实例数据集和有效的空间解耦机制。为此,我们构建了MILD数据集,涵盖多样姿态、遮挡关系与复杂多实例交互。定义跨域注意力差距(CAG)作为语义泄漏量化指标,并提出多层扩散(MILD)策略,将生成过程分解为独立去噪路径,实现各前景实例与背景的分别重建。为增强人体中心理解,引入人体形态引导模块,融合姿态、分割与空间关系信息以提升结构感知能力。此外,提出空间调制注意力机制,利用空间掩码先验动态调节语义区域注意力,进一步扩大CAG以减少边界伪影和语义泄漏。实验表明,MILD显著优于现有方法。数据集与代码已公开:https://mild-multi-layer-diffusion.github.io/。
原文摘要 · Abstract (English)
Recent years have witnessed the success of diffusion models in image customization tasks. However, existing mask-guided human erasing methods still struggle in complex scenarios such as human-human occlusion, human-object entanglement, and human-background interference, mainly due to the lack of large-scale multi-instance datasets and effective spatial decoupling to separate foreground from background. To bridge these gaps, we curate the MILD dataset capturing diverse poses, occlusions, and complex multi-instance interactions. We then define the Cross-Domain Attention Gap (CAG), an attention-gap metric to quantify semantic leakage. On top of these, we propose Multi-Layer Diffusion (MILD), which decomposes the generation process into independent denoising pathways, enabling separate reconstruction of each foreground instance and the background. To enhance human-centric understanding, we introduce Human Morphology Guidance, a plug-and-play module that incorporates pose, parsing, and spatial relationships into the diffusion process to improve structural awareness and restoration quality. Additionally, we present Spatially-Modulated Attention, an adaptive mechanism that leverages spatial mask priors to modulate attention across semantic regions, further widening the CAG to effectively minimize boundary artifacts and mitigate semantic leakage. Experiments show that MILD significantly outperforms existing methods. Datasets and code are publicly available at: https://mild-multi-layer-diffusion.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。