统一框架实现多模态线稿上色,支持任意组合控制信号。
OmniColor: A Unified Framework for Multi-modal Lineart Colorization
- 分两类控制信号:空间对齐与语义参考,分别用双路径编码和VLM-only方案处理。
- 引入自适应门控模块解决多模态冲突,实现边界精准保留与颜色高保真还原。
- 适用于需要灵活可控上色的专业创作,尤其适合影视动画流程。
线稿上色是专业内容创作的关键环节,但在多样用户约束下实现精确且灵活的结果仍具挑战。为此,我们提出OmniColor,一种支持任意组合控制信号的多模态线稿上色统一框架。我们将引导信号系统性地分为两类:空间对齐条件与语义参考条件。针对空间对齐输入,采用双路径编码策略并结合密集特征对齐损失,确保严格边界保持与精确色彩恢复;针对语义参考输入,使用仅基于视觉语言模型(VLM)的编码方案,并集成时间冗余消除机制,过滤重复信息以提升推理效率。为解决潜在输入冲突,引入自适应空间-语义门控模块,动态平衡多模态约束。实验表明,OmniColor在可控性、视觉质量与时间稳定性方面均表现优异,提供了一种鲁棒且实用的线稿上色解决方案。源代码与数据集将开源于https://github.com/zhangxulu1996/OmniColor。
原文摘要 · Abstract (English)
Lineart colorization is a critical stage in professional content creation, yet achieving precise and flexible results under diverse user constraints remains a significant challenge. To address this, we propose OmniColor, a unified framework for multi-modal lineart colorization that supports arbitrary combinations of control signals. Specifically, we systematically categorize guidance signals into two types: spatially-aligned conditions and semantic-reference conditions. For spatially-aligned inputs, we employ a dual-path encoding strategy paired with a Dense Feature Alignment loss to ensure rigorous boundary preservation and precise color restoration. For semantic-reference inputs, we utilize a VLM-only encoding scheme integrated with a Temporal Redundancy Elimination mechanism to filter repetitive information and enhance inference efficiency. To resolve potential input conflicts, we introduce an Adaptive Spatial-Semantic Gating module that dynamically balances multi-modal constraints. Experimental results demonstrate that OmniColor achieves superior controllability, visual quality, and temporal stability, providing a robust and practical solution for lineart colorization. The source code and dataset will be open at https://github.com/zhangxulu1996/OmniColor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。