arXiv:2508.09131cs.GRcs.AI2025-08被引 4

无需训练即可精准控制图像颜色,保持物理一致性。

Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer

  • 利用多模态扩散模型的注意力机制分离结构与颜色,精准编辑。
  • 在SD3和FLUX.1-dev上优于现有方法,一致性和质量达顶尖水平。
  • 适合需要精细色彩控制的视觉生成、视频编辑场景。

图像与视频中的文本引导颜色编辑是一个基础但尚未解决的问题,需对反照率、光源颜色和环境光照等颜色属性进行细粒度操控,同时保持几何结构、材质特性及光-物质相互作用的物理一致性。现有无训练方法虽适用范围广,但在颜色控制精度和编辑区域一致性方面表现不佳。本文提出ColorCtrl,一种基于现代多模态扩散变换器(MM-DiT)注意力机制的无训练颜色编辑方法。通过针对性地操作注意力图与值令牌,实现结构与颜色的解耦,支持词级属性强度控制。仅修改提示指定区域,不影响其他部分。在SD3和FLUX.1-dev上的大量实验表明,ColorCtrl超越现有无训练方法,达到当前最佳的编辑质量与一致性。此外,其表现优于FLUX.1 Kontext Max和GPT-4o Image Generation在一致性方面的表现。扩展至CogVideoX视频模型时,该方法在时间连贯性和编辑稳定性上优势更明显。同时,其也适用于Step1X-Edit和FLUX.1 Kontext dev等指令式编辑扩散模型,展现出强泛化能力。

原文摘要 · Abstract (English)

Text-guided color editing in images and videos is a fundamental yet unsolved problem, requiring fine-grained manipulation of color attributes, including albedo, light source color, and ambient lighting, while preserving physical consistency in geometry, material properties, and light-matter interactions. Existing training-free methods offer broad applicability across editing tasks but struggle with precise color control and often introduce visual inconsistency in both edited and non-edited regions. In this work, we present ColorCtrl, a training-free color editing method that leverages the attention mechanisms of modern Multi-Modal Diffusion Transformers (MM-DiT). By disentangling structure and color through targeted manipulation of attention maps and value tokens, our method enables accurate and consistent color editing, along with word-level control of attribute intensity. Our method modifies only the intended regions specified by the prompt, leaving unrelated areas untouched. Extensive experiments on both SD3 and FLUX.1-dev demonstrate that ColorCtrl outperforms existing training-free approaches and achieves state-of-the-art performances in both edit quality and consistency. Furthermore, our method surpasses strong commercial models such as FLUX.1 Kontext Max and GPT-4o Image Generation in terms of consistency. When extended to video models like CogVideoX, our approach exhibits greater advantages, particularly in maintaining temporal coherence and editing stability. Finally, our method also generalizes to instruction-based editing diffusion models such as Step1X-Edit and FLUX.1 Kontext dev, further demonstrating its versatility.

颜色编辑扩散模型无训练视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。