用文字提示提升黑白图上色质量,效果可量化
Multimodal Image Colorization: Quantifying the Impact of Text-Conditioned Guidance on Grayscale-to-Color Translation

- 对比有无文本条件的U-Net与Stable Diffusion模型
- 文本引导使色彩准确度提升36.6%,视觉相似性提高1.5%
- 适合图像修复、艺术创作等需精准上色的场景
灰度图像广泛存在于历史摄影修复、医学成像和艺术媒体中。然而,自动为其上色仍是计算机视觉中的重大挑战,因为同一灰度输入可能存在多种合理着色方案。本文量化了文本条件对灰度到彩色图像转换模型在像素级和感知指标上的影响。具体比较了U-Net与Stable Diffusion 1.5两种架构,分别在有无CLIP文本条件的情况下进行测试,其余变量保持一致。结果显示,在U-Net中,文本条件使PSNR提升5.6%,SSIM提升1.2%,色彩丰富度提升36.6%,LPIPS降低7.6%;在Stable Diffusion中,相应指标分别为PSNR提升5.8%,SSIM提升1.5%,色彩丰富度提升0.6%,LPIPS降低11.3%。结果表明,文本条件在不同架构下均带来稳定且可量化的着色质量提升。
原文摘要 · Abstract (English)
Grayscale images are commonly found in historical photography restoration, medical imaging, and artistic media. However, automatically applying color to these images remains a significant challenge in computer vision because many plausible colorizations can correspond to the same grayscale input. In this work, we quantify the effect of text conditioning on pixel-level and perceptual metrics for grayscale-to-color image models. Specifically, we compare two architectures, a U-Net and Stable Diffusion 1.5, each tested with and without CLIP text conditioning while holding all other variables constant. Our results show that text conditioning improves PSNR by 5.6%, SSIM by 1.2%, and colorfulness by 36.6%, while reducing LPIPS by 7.6% in the U-Net tier. In the Stable Diffusion tier, text conditioning improves PSNR by 5.8%, SSIM by 1.5%, and colorfulness by 0.6%, while reducing LPIPS by 11.3%. These results indicate that text conditioning provides consistent, measurable improvements to colorization quality across both architecture scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。