仅用两张图训练,实现老照片精准上色且不扭曲结构
Structure-preserving Feature Alignment for Old Photo Colorization
- 通过特征分布对齐建立参考图与老照片的语义对应关系
- 在仅两张图训练下实现媲美大数据模型的上色效果
- 引入感知约束和金字塔结构保持机制,避免颜色迁移导致形变
深度学习在基于参考图像的上色任务中取得显著进展,但直接应用于老照片上色时面临缺乏真实标签和自然灰度图与老照片之间存在显著领域差异的问题。为此,本文提出一种新型基于CNN的算法SFAC(Structure-preserving Feature Alignment Colorizer),仅需两张图像即可训练,摆脱对大规模数据的依赖,并可直接处理老照片以缓解领域差距。核心目标是建立两图间的语义对应关系,使语义相关物体具有相似颜色。通过特征分布对齐损失实现对不同度量方式的鲁棒性。然而,仅依赖语义对应进行颜色迁移可能导致结构失真。为此,我们引入结构保持机制,在特征层面加入感知约束,在像素层面采用冻结-更新金字塔结构。大量实验表明,该方法在定性和定量指标上均表现优异,有效提升了老照片上色质量。
原文摘要 · Abstract (English)
Deep learning techniques have made significant advancements in reference-based colorization by training on large-scale datasets. However, directly applying these methods to the task of colorizing old photos is challenging due to the lack of ground truth and the notorious domain gap between natural gray images and old photos. To address this issue, we propose a novel CNN-based algorithm called SFAC, i.e., Structure-preserving Feature Alignment Colorizer. SFAC is trained on only two images for old photo colorization, eliminating the reliance on big data and allowing direct processing of the old photo itself to overcome the domain gap problem. Our primary objective is to establish semantic correspondence between the two images, ensuring that semantically related objects have similar colors. We achieve this through a feature distribution alignment loss that remains robust to different metric choices. However, utilizing robust semantic correspondence to transfer color from the reference to the old photo can result in inevitable structure distortions. To mitigate this, we introduce a structure-preserving mechanism that incorporates a perceptual constraint at the feature level and a frozen-updated pyramid at the pixel level. Extensive experiments demonstrate the effectiveness of our method for old photo colorization, as confirmed by qualitative and quantitative metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。