arXiv:2410.04972cs.CV2024-10被引 3

用文字描述给黑白视频上色,既自由又不闪动。

L-C4: Language-Based Video Colorization for Creative and Consistent Color

  • 用用户写的文字描述控制上色,突破传统检索限制。
  • 在多个数据集上优于现有方法,颜色语义准确且连续稳定。
  • 适合需要创意配色和长时间一致性的视频修复场景。

自动视频上色本质上是病态问题,因为每帧灰度图像都有多种可能的色彩选项。以往基于样例的视频上色方法因复杂的样例检索过程限制了用户的想象力;而基于条件的图像上色结合后处理算法仍难以保持时间一致性。为此,我们提出语言引导的视频上色方法L-C4,通过用户提供的文字描述指导上色过程。模型基于预训练的跨模态生成模型,利用其强大的语言理解与色彩表征能力。引入跨模态预融合模块生成实例感知的文本嵌入,实现创意色彩应用;提出时序可变形注意力机制防止闪烁或颜色漂移,并设计跨裁剪融合策略维持长时色彩一致性。大量实验表明,L-C4在多个基准上超越现有方法,实现了语义准确的颜色、无约束的创意匹配以及强鲁棒的时间一致性。

原文摘要 · Abstract (English)

Automatic video colorization is inherently an ill-posed problem because each monochrome frame has multiple optional color candidates. Previous exemplar-based video colorization methods restrict the user's imagination due to the elaborate retrieval process. Alternatively, conditional image colorization methods combined with post-processing algorithms still struggle to maintain temporal consistency. To address these issues, we present Language-based video Colorization for Creative and Consistent Colors (L-C4) to guide the colorization process using user-provided language descriptions. Our model is built upon a pre-trained cross-modality generative model, leveraging its comprehensive language understanding and robust color representation abilities. We introduce the cross-modality pre-fusion module to generate instance-aware text embeddings, enabling the application of creative colors. Additionally, we propose temporally deformable attention to prevent flickering or color shifts, and cross-clip fusion to maintain long-term color consistency. Extensive experimental results demonstrate that L-C4 outperforms relevant methods, achieving semantically accurate colors, unrestricted creative correspondence, and temporally robust consistency.

视频上色文本生成跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。