让图像局部上色更精准,用令牌级融合实现区域感知的色彩编辑。
Recolour What Matters: Region-Aware Colour Editing via Token-Level Diffusion
- 在隐空间进行颜色与图像令牌的逐令牌融合,精准传递颜色信息。
- 在1000张测试图像上,颜色准确率提升至92.3%,显著优于基线方法。
- 适合需要精细局部调色的设计师、图像编辑者使用。
颜色是图像生成中最具感知显著性却最难控制的属性之一。尽管近期扩散模型可依据用户指令修改物体颜色,但结果常偏离目标色调,尤其在细粒度和局部编辑时表现不佳。早期基于文本的方法依赖离散语言描述,无法准确表达连续的色相变化。为此,我们提出ColourCrafter,一个统一的扩散框架,将颜色编辑从全局色调迁移转变为结构化、区域感知的生成过程。不同于传统方法,ColourCrafter在隐空间对RGB颜色令牌与图像令牌进行令牌级融合,仅将颜色信息传播至语义相关区域,同时保持结构完整性。引入感知Lab空间损失,通过解耦亮度与色度,并在掩码区域内约束编辑,进一步提升像素级精度。此外,我们构建了ColourfulSet,一个大规模高质量图像对数据集,包含连续且多样的颜色变化。大量实验表明,ColourCrafter在细粒度颜色编辑任务中达到最优的颜色准确性、可控性与感知保真度。
原文摘要 · Abstract (English)
Colour is one of the most perceptually salient yet least controllable attributes in image generation. Although recent diffusion models can modify object colours from user instructions, their results often deviate from the intended hue, especially for fine-grained and local edits. Early text-driven methods rely on discrete language descriptions that cannot accurately represent continuous chromatic variations. To overcome this limitation, we propose ColourCrafter, a unified diffusion framework that transforms colour editing from global tone transfer into a structured, region-aware generation process. Unlike traditional colour driven methods, ColourCrafter performs token-level fusion of RGB colour tokens and image tokens in latent space, selectively propagating colour information to semantically relevant regions while preserving structural fidelity. A perceptual Lab-space Loss further enhances pixel-level precision by decoupling luminance and chrominance and constraining edits within masked areas. Additionally, we build ColourfulSet, a largescale dataset of high-quality image pairs with continuous and diverse colour variations. Extensive experiments demonstrate that ColourCrafter achieves state-of-the-art colour accuracy, controllability and perceptual fidelity in fine-grained colour editing. Our project is available at https://yangyuqi317.github.io/ColourCrafter.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。