用单阶段扩散模型实现高保真人像超分,修复面部更自然。
HeadsUp! High-Fidelity Portrait Image Super-Resolution
- 基于单步扩散模型,端到端完成人像超分。
- 引入人脸监督与参考机制,提升面部细节与身份一致性。
- 自建4K人像数据集,适合追求真实感的图像生成研究者。
人像照片在社交媒体中极为常见,通常包含人物主体与自然背景。现有图像超分辨率(ISR)技术多聚焦于通用图像或严格对齐的人脸图像(即人脸超分)。实际应用中,常通过组合专用人脸模型与通用模型处理人像:前者负责面部区域,后者处理其余部分。然而,这种拼接方法因训练方式差异,不可避免在面部区域引入融合或边界伪影,而人类对人脸保真度极为敏感。为克服此问题,我们研究人像超分(PortraitISR)任务,提出一种单阶段扩散模型 HeadsUp,可端到端无缝恢复与放大人像图像。具体而言,我们在单步扩散模型基础上,设计人脸监督机制以引导模型关注面部区域;并引入基于参考的机制,辅助身份还原,降低低质量人脸恢复中的模糊性。此外,我们构建了高质量的4K人像超分数据集 PortraitSR-4K,用于模型训练与基准测试。大量实验表明,HeadsUp 在 PortraitISR 任务上达到当前最优性能,同时在通用图像与对齐人脸数据集上保持相当或更高的表现。
原文摘要 · Abstract (English)
Portrait pictures, which typically feature both human subjects and natural backgrounds, are one of the most prevalent forms of photography on social media. Existing image super-resolution (ISR) techniques generally focus either on generic real-world images or strictly aligned facial images (i.e., face super-resolution). In practice, separate models are blended to handle portrait photos: the face specialist model handles the face region, and the general model processes the rest. However, these blending approaches inevitably introduce blending or boundary artifacts around the facial regions due to different model training recipes, while human perception is particularly sensitive to facial fidelity. To overcome these limitations, we study the portrait image supersolution (PortraitISR) problem, and propose HeadsUp, a single-step diffusion model that is capable of seamlessly restoring and upscaling portrait images in an end-to-end manner. Specifically, we build our model on top of a single-step diffusion model and develop a face supervision mechanism to guide the model in focusing on the facial region. We then integrate a reference-based mechanism to help with identity restoration, reducing face ambiguity in low-quality face restoration. Additionally, we have built a high-quality 4K portrait image ISR dataset dubbed PortraitSR-4K, to support model training and benchmarking for portrait images. Extensive experiments show that HeadsUp achieves state-of-the-art performance on the PortraitISR task while maintaining comparable or higher performance on both general image and aligned face datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。