通过分组件注意力与坐标保持融合,提升手绘草图转真实图像的细节还原与空间一致性。
Component-Aware Sketch-to-Image Generation Using Self-Attention Encoding and Coordinate-Preserving Fusion
- 分组件建模+自注意力编码,精准捕捉草图局部语义结构
- 在CelebAMask-HQ上比现有方法提升21% FID、58% IS等指标
- 适合需要高保真草图生成的数字艺术修复与刑侦图像重建
将自由手绘草图转化为逼真图像仍是图像合成中的核心挑战,尤其因草图具有抽象、稀疏和风格多样的特点。现有基于GAN和扩散模型的方法常难以重建细粒度细节、保持空间对齐或跨领域适应。本文提出一种分组件感知、自修正的两阶段草图到图像生成框架。首先,自注意力编码器(SA2N)从草图各组件区域提取局部语义与结构特征;随后,坐标保持门控融合模块(CGF)将其整合为连贯的空间布局;最后,基于改进StyleGAN2的时空自适应精修模块(SARR)通过迭代优化增强真实感与一致性。在人脸(CelebAMask-HQ、CUFSF)与非人脸(Sketchy、ChairsV2、ShoesV2)数据集上的实验表明,该方法具备强鲁棒性与泛化能力。相比顶尖GAN与扩散模型,在图像保真度、语义准确性与感知质量上均显著领先:在CelebAMask-HQ上实现FID降低21%、IS提升58%、KID下降41%、SSIM提升20%。结合更高效率与跨域视觉一致性,本方法在刑侦、数字艺术修复及通用草图图像合成中具有应用潜力。
原文摘要 · Abstract (English)
Translating freehand sketches into photorealistic images remains a fundamental challenge in image synthesis, particularly due to the abstract, sparse, and stylistically diverse nature of sketches. Existing approaches, including GAN-based and diffusion-based models, often struggle to reconstruct fine-grained details, maintain spatial alignment, or adapt across different sketch domains. In this paper, we propose a component-aware, self-refining framework for sketch-to-image generation that addresses these challenges through a novel two-stage architecture. A Self-Attention-based Autoencoder Network (SA2N) first captures localised semantic and structural features from component-wise sketch regions, while a Coordinate-Preserving Gated Fusion (CGF) module integrates these into a coherent spatial layout. Finally, a Spatially Adaptive Refinement Revisor (SARR), built on a modified StyleGAN2 backbone, enhances realism and consistency through iterative refinement guided by spatial context. Extensive experiments across both facial (CelebAMask-HQ, CUFSF) and non-facial (Sketchy, ChairsV2, ShoesV2) datasets demonstrate the robustness and generalizability of our method. The proposed framework consistently outperforms state-of-the-art GAN and diffusion models, achieving significant gains in image fidelity, semantic accuracy, and perceptual quality. On CelebAMask-HQ, our model improves over prior methods by 21% (FID), 58% (IS), 41% (KID), and 20% (SSIM). These results, along with higher efficiency and visual coherence across diverse domains, position our approach as a strong candidate for applications in forensics, digital art restoration, and general sketch-based image synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。