arXiv:2409.03644cs.CV2024-09AAAI被引 12

用两阶段方法修复生成图像中畸形的人体部位,提升真实感。

RealisHuman: A Two-Stage Approach for Refining Malformed Human Parts in Generated Images

  • 先参考错误部位生成真实人体部件,再局部重绘融合
  • 修复后图像在视觉和量化指标上均有显著提升
  • 适合需要高质量人像生成的研究者与开发者

近年来,扩散模型已颠覆视觉生成领域,超越传统生成对抗网络(GANs)。然而,由于人体结构复杂,生成具有真实语义部件(如手部、面部)的图像仍面临重大挑战。为此,我们提出一种新型后处理方案 RealisHuman。该框架分两阶段运行:第一阶段以原始畸形部位为参考,生成更真实的头部或手部等人体部件,确保与原图细节一致;第二阶段通过重绘周围区域,将修复后的部件无缝嵌入原位置,实现平滑自然的融合。RealisHuman 显著提升了人像生成的真实感,在定性和定量评估中均表现优异。代码已公开于 https://github.com/Wangbenzhi/RealisHuman。

原文摘要 · Abstract (English)

In recent years, diffusion models have revolutionized visual generation, outperforming traditional frameworks like Generative Adversarial Networks (GANs). However, generating images of humans with realistic semantic parts, such as hands and faces, remains a significant challenge due to their intricate structural complexity. To address this issue, we propose a novel post-processing solution named RealisHuman. The RealisHuman framework operates in two stages. First, it generates realistic human parts, such as hands or faces, using the original malformed parts as references, ensuring consistent details with the original image. Second, it seamlessly integrates the rectified human parts back into their corresponding positions by repainting the surrounding areas to ensure smooth and realistic blending. The RealisHuman framework significantly enhances the realism of human generation, as demonstrated by notable improvements in both qualitative and quantitative metrics. Code is available at https://github.com/Wangbenzhi/RealisHuman.

图像修复扩散模型人体生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。