用噪声引导的注意力机制,让素描生成逼真人脸更精准。
Locally-Focused Face Representation for Sketch-to-Image Generation Using Noise-Induced Refinement
- 通过注意力模块聚焦关键面部特征,提升细节还原度
- 在三大数据集上FID分别降低17、23、38,效果领先
- 可跨类型素描生成,适合刑侦等实际应用
本文提出一种新型深度学习框架,显著提升素描到真实人脸图像的转换质量。采用基于卷积块注意力的自编码网络(CA2N),在编码器-解码器结构中通过块注意力机制有效捕捉并增强关键面部特征。随后,利用噪声诱导的条件生成对抗网络(cGAN)进行细化,使系统在未见领域仍保持高性能。该方法大幅提高图像真实感与保真度,在CelebAMask-HQ、CUHK、CUFSF数据集上,模型的FID分别比最优现有方法降低17、23、38,达到该任务新基准。模型具备良好泛化能力,可处理多种素描类型,适用于刑事侦查等实际场景。
原文摘要 · Abstract (English)
This paper presents a novel deep-learning framework that significantly enhances the transformation of rudimentary face sketches into high-fidelity colour images. Employing a Convolutional Block Attention-based Auto-encoder Network (CA2N), our approach effectively captures and enhances critical facial features through a block attention mechanism within an encoder-decoder architecture. Subsequently, the framework utilises a noise-induced conditional Generative Adversarial Network (cGAN) process that allows the system to maintain high performance even on domains unseen during the training. These enhancements lead to considerable improvements in image realism and fidelity, with our model achieving superior performance metrics that outperform the best method by FID margin of 17, 23, and 38 on CelebAMask-HQ, CUHK, and CUFSF datasets; respectively. The model sets a new state-of-the-art in sketch-to-image generation, can generalize across sketch types, and offers a robust solution for applications such as criminal identification in law enforcement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。