arXiv:2602.23191cs.CV2026-02

统一图像与视频素描上色,细节更准、动态更稳。

Uni-Animator: Towards Unified Visual Colorization

  • 用实例块嵌入增强参考颜色对齐,提升着色精度。
  • 引入物理特征强化高频纹理,保留更多细节。
  • 动态RoPE编码建模运动依赖,减少视频抖动伪影。

我们提出Uni-Animator,一种基于扩散Transformer(DiT)的统一图像与视频素描上色框架。现有方法难以兼顾图像与视频任务,存在单/多参考下着色不准、高频物理细节保留不足、大运动场景中时间不一致与运动伪影等问题。为解决着色不准问题,我们引入实例块嵌入进行视觉参考增强,实现参考颜色信息的精准对齐与融合;为改善细节丢失,设计基于物理特征的细节强化机制,有效捕捉并保留高频纹理;为缓解运动导致的时间不一致性,提出基于素描的动态RoPE编码,自适应建模时空依赖关系。大量实验表明,Uni-Animator在图像与视频素描上色任务上均达到与专用方法相当的性能,兼具高细节保真度与强时间一致性,实现跨域统一能力。

原文摘要 · Abstract (English)

We propose Uni-Animator, a novel Diffusion Transformer (DiT)-based framework for unified image and video sketch colorization. Existing sketch colorization methods struggle to unify image and video tasks, suffering from imprecise color transfer with single or multiple references, inadequate preservation of high-frequency physical details, and compromised temporal coherence with motion artifacts in large-motion scenes. To tackle imprecise color transfer, we introduce visual reference enhancement via instance patch embedding, enabling precise alignment and fusion of reference color information. To resolve insufficient physical detail preservation, we design physical detail reinforcement using physical features that effectively capture and retain high-frequency textures. To mitigate motion-induced temporal inconsistency, we propose sketch-based dynamic RoPE encoding that adaptively models motion-aware spatial-temporal dependencies. Extensive experimental results demonstrate that Uni-Animator achieves competitive performance on both image and video sketch colorization, matching that of task-specific methods while unlocking unified cross-domain capabilities with high detail fidelity and robust temporal consistency.

图像生成视频生成扩散模型上色

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。