用文字描述精准还原黑白图色彩,速度超旧方法14倍
Language-based Image Colorization: A Benchmark and Beyond
- 基于蒸馏扩散模型,通过文本引导实现跨模态对齐
- 在多个数据集上超越复杂模型,速度提升14倍
- 首个系统性评测语言驱动图像着色的综述与基准
图像着色旨在为灰度图像恢复颜色。自动着色方法因颜色歧义问题难以生成高质量结果,且用户控制能力有限。随着跨模态数据集和模型的发展,语言驱动着色方法利用文本描述灵活引导着色过程。针对该领域缺乏全面综述的问题,本文开展系统分析与基准测试。首先简要总结现有自动着色方法;重点分析语言驱动方法的核心挑战——跨模态对齐。将现有方法分为两类:一类从头训练跨模态网络,另一类利用预训练跨模态模型建立图文对应关系。基于分析局限性,提出一种简单而有效的基于蒸馏扩散模型的方法。大量实验表明,该简单基线在性能上优于以往复杂方法,且速度提升14倍。据我们所知,这是首个关于语言驱动图像着色领域的综合性综述与基准,为社区提供重要洞察。代码已开源。
原文摘要 · Abstract (English)
Image colorization aims to bring colors back to grayscale images. Automatic image colorization methods, which requires no additional guidance, struggle to generate high-quality images due to color ambiguity, and provides limited user controllability. Thanks to the emergency of cross-modality datasets and models, language-based colorization methods are proposed to fully utilize the efficiency and flexibly of text descriptions to guide colorization. In view of the lack of a comprehensive review of language-based colorization literature, we conduct a thorough analysis and benchmarking. We first briefly summarize existing automatic colorization methods. Then, we focus on language-based methods and point out their core challenge on cross-modal alignment. We further divide these methods into two categories: one attempts to train a cross-modality network from scratch, while the other utilizes the pre-trained cross-modality model to establish the textual-visual correspondence. Based on the analyzed limitations of existing language-based methods, we propose a simple yet effective method based on distilled diffusion model. Extensive experiments demonstrate that our simple baseline can produces better results than previous complex methods with 14 times speed up. To the best of our knowledge, this is the first comprehensive review and benchmark on language-based image colorization field, providing meaningful insights for the community. The code is available at https://github.com/lyf1212/Color-Turbo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。