arXiv:2506.04283cs.GRcs.AI2025-06NeurIPS

用SSIM引导的扩散模型,让动漫人脸草图自动上色更准更自然。

SSIMBaD: Sigma Scaling with SSIM-Guided Balanced Diffusion for AnimeFace Colorization

  • 通过SSIM度量结构相似性,动态调整噪声尺度以平衡各步骤难度。
  • 在动漫人脸数据集上,像素准确率和视觉质量均优于现有方法。
  • 适合需要高质量动漫风格上色的创作者或工业应用。

我们提出一种基于扩散模型的动漫人脸草图自动上色新框架。该方法在保持输入草图结构完整性的同时,有效迁移参考图像的风格特征。与依赖预设噪声调度的传统方法不同,本框架基于连续时间扩散模型,引入SSIMBaD(Sigma Scaling with SSIM-Guided Balanced Diffusion)。SSIMBaD采用sigma空间变换,使结构相似性(SSIM)衡量的感知退化呈线性变化,确保各时间步视觉重建难度一致,从而实现更均衡、更忠实的还原。在大规模动漫人脸数据集上的实验表明,该方法在像素精度和感知质量上均超越当前最优模型,且对多种风格具有良好的泛化能力。代码已开源。

原文摘要 · Abstract (English)

We propose a novel diffusion-based framework for automatic colorization of Anime-style facial sketches. Our method preserves the structural fidelity of the input sketch while effectively transferring stylistic attributes from a reference image. Unlike traditional approaches that rely on predefined noise schedules - which often compromise perceptual consistency -- our framework builds on continuous-time diffusion models and introduces SSIMBaD (Sigma Scaling with SSIM-Guided Balanced Diffusion). SSIMBaD applies a sigma-space transformation that aligns perceptual degradation, as measured by structural similarity (SSIM), in a linear manner. This scaling ensures uniform visual difficulty across timesteps, enabling more balanced and faithful reconstructions. Experiments on a large-scale Anime face dataset demonstrate that our method outperforms state-of-the-art models in both pixel accuracy and perceptual quality, while generalizing to diverse styles. Code is available at github.com/Giventicket/SSIMBaD-Sigma-Scaling-with-SSIM-Guided-Balanced-Diffusion-for-AnimeFace-Colorization

动漫上色扩散模型结构相似性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。