arXiv:2511.22330cs.CVcs.AI2025-11

用语言提示自动给视频上色,无闪烁且效果逼真。

Prompt-based Consistent Video Colorization

  • 用语言和分割掩码引导扩散模型上色
  • 在DAVIS30和VIDEVO20上达到最高PSNR与视觉真实度
  • 自动对齐前后帧颜色,修复失真

现有视频上色方法存在时间闪烁问题或需大量人工输入。本文提出一种新方法,利用语言与分割提供的丰富语义引导,实现高保真自动上色。采用语言条件扩散模型对灰度帧上色,通过自动生成的物体掩码和文本提示提供指导;主要自动方法使用通用提示,在无需特定颜色输入的情况下达到当前最佳性能。通过光流(RAFT)将前一帧颜色信息逐帧对齐以保证时序稳定性,并设计修正步骤检测并修复对齐引入的不一致性。在DAVIS30和VIDEVO20标准基准上评估显示,该方法在色彩准确性(PSNR)与视觉真实感(Colorfulness, CDC)方面均达最优,验证了基于自动提示的引导在一致视频上色中的有效性。

原文摘要 · Abstract (English)

Existing video colorization methods struggle with temporal flickering or demand extensive manual input. We propose a novel approach automating high-fidelity video colorization using rich semantic guidance derived from language and segmentation. We employ a language-conditioned diffusion model to colorize grayscale frames. Guidance is provided via automatically generated object masks and textual prompts; our primary automatic method uses a generic prompt, achieving state-of-the-art results without specific color input. Temporal stability is achieved by warping color information from previous frames using optical flow (RAFT); a correction step detects and fixes inconsistencies introduced by warping. Evaluations on standard benchmarks (DAVIS30, VIDEVO20) show our method achieves state-of-the-art performance in colorization accuracy (PSNR) and visual realism (Colorfulness, CDC), demonstrating the efficacy of automated prompt-based guidance for consistent video colorization.

视频上色扩散模型语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。