arXiv:2501.09499cs.CV2025-01被引 2

用扩散模型实现视频着色,解决颜色溢出和闪烁问题。

VanGogh: A Unified Multimodal Diffusion-based Framework for Video Colorization

  • 多模态融合+深度引导生成,减少颜色溢出。
  • 支持全局与局部控制,提升时序一致性与色彩保真度。
  • 适合需要高质量视频着色的影视修复与内容创作场景。

视频着色旨在将灰度视频转换为生动的彩色表示,同时保持时间一致性和结构完整性。现有方法常出现颜色溢出问题,且在复杂运动或多样语义提示下缺乏全面控制。为此,我们提出VanGogh,一种统一的多模态扩散框架用于视频着色。该框架采用双Qformer对齐并融合多模态特征,结合深度引导生成过程与光流损失,有效减少颜色溢出。此外,引入颜色注入策略和亮度通道替换,提升泛化能力并缓解闪烁伪影。该设计使用户可实现全局与局部控制,生成更高质量的彩色视频。大量定性与定量评估及用户研究显示,VanGogh在时间一致性与色彩保真度方面表现优异。

原文摘要 · Abstract (English)

Video colorization aims to transform grayscale videos into vivid color representations while maintaining temporal consistency and structural integrity. Existing video colorization methods often suffer from color bleeding and lack comprehensive control, particularly under complex motion or diverse semantic cues. To this end, we introduce VanGogh, a unified multimodal diffusion-based framework for video colorization. VanGogh tackles these challenges using a Dual Qformer to align and fuse features from multiple modalities, complemented by a depth-guided generation process and an optical flow loss, which help reduce color overflow. Additionally, a color injection strategy and luma channel replacement are implemented to improve generalization and mitigate flickering artifacts. Thanks to this design, users can exercise both global and local control over the generation process, resulting in higher-quality colorized videos. Extensive qualitative and quantitative evaluations, and user studies, demonstrate that VanGogh achieves superior temporal consistency and color fidelity.Project page: https://becauseimbatman0.github.io/VanGogh.

视频着色扩散模型多模态时序一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。