arXiv:2508.09709cs.CV2025-08被引 2

用注意力机制自动对齐漫画线稿与参考图色彩,解决姿势不同时的配色不一致问题。

MangaDiT: Reference-Guided Line Art Colorization with Hierarchical Attention in Diffusion Transformers

  • 通过分层注意力和动态加权,隐式发现参考图与目标图的语义对应关系。
  • 在两个基准数据集上均超越现有方法,实现更一致的区域色彩匹配。
  • 适合需要高质量漫画自动着色的创作者或自动化内容生产团队。

扩散模型的进展显著提升了参考图引导的漫画线稿着色效果。然而,现有方法在角色姿态或动作不同时仍存在区域级配色不一致的问题。本文提出MangaDiT,一种基于扩散Transformer(DiT)的参考图引导线稿着色模型,无需外部匹配标注,而是通过内部注意力机制隐式发现语义对应。该模型引入分层注意力机制与动态加权策略,增加一个上下文感知路径,利用池化空间特征扩展感受野,增强区域级色彩对齐能力。在两个基准数据集上的实验表明,本方法在定性与定量评估中均显著优于当前最优方法。

原文摘要 · Abstract (English)

Recent advances in diffusion models have significantly improved the performance of reference-guided line art colorization. However, existing methods still struggle with region-level color consistency, especially when the reference and target images differ in character pose or motion. Instead of relying on external matching annotations between the reference and target, we propose to discover semantic correspondences implicitly through internal attention mechanisms. In this paper, we present MangaDiT, a powerful model for reference-guided line art colorization based on Diffusion Transformers (DiT). Our model takes both line art and reference images as conditional inputs and introduces a hierarchical attention mechanism with a dynamic attention weighting strategy. This mechanism augments the vanilla attention with an additional context-aware path that leverages pooled spatial features, effectively expanding the model's receptive field and enhancing region-level color alignment. Experiments on two benchmark datasets demonstrate that our method significantly outperforms state-of-the-art approaches, achieving superior performance in both qualitative and quantitative evaluations.

漫画着色扩散模型注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。