用文字或参考图精准控制图像调色,让色彩更符合审美意图。
AceTone: Bridging Words and Colors for Conditional Image Grading
- 基于文本或图片生成3D-LUT调色参数,实现统一框架下的条件调色
- 在80万张图像数据上训练,相比现有方法提升50%的感知相似度
- 支持创意风格一致性和人类审美偏好,适合影视与设计场景
色彩影响图像风格与情感表达。以往调色方法依赖局部重着色或固定滤波器,难以适应多样创作意图或匹配人类审美。本文提出AceTone,首个支持多模态条件调色的统一框架。将调色建模为生成式颜色变换任务,模型根据文本提示或参考图像直接生成3D-LUT。设计基于VQ-VAE的编码器,将3×32³大小的LUT向量压缩为64个离散码本项,保持ΔE<2的保真度。构建大规模数据集AceTone-800K,训练视觉-语言模型预测LUT码项,并通过强化学习优化输出的感知保真度与美学效果。实验表明,AceTone在文本引导与参考图像引导调色任务中均达到领先性能,LPIPS指标最高提升50%。人工评估确认其结果视觉悦目且风格连贯,开辟了语言驱动、美学对齐调色的新路径。
原文摘要 · Abstract (English)
Color affects how we interpret image style and emotion. Previous color grading methods rely on patch-wise recoloring or fixed filter banks, struggling to generalize across creative intents or align with human aesthetic preferences. In this study, we propose AceTone, the first approach that supports multimodal conditioned color grading within a unified framework. AceTone formulates grading as a generative color transformation task, where a model directly produces 3D-LUTs conditioned on text prompts or reference images. We develop a VQ-VAE based tokenizer which compresses a $3\times32^3$ LUT vector to 64 discrete tokens with $ΔE<2$ fidelity. We further build a large-scale dataset, AceTone-800K, and train a vision-language model to predict LUT tokens, followed by reinforcement learning to align outputs with perceptual fidelity and aesthetics. Experiments show that AceTone achieves state-of-the-art performance on both text-guided and reference-guided grading tasks, improving LPIPS by up to 50% over existing methods. Human evaluations confirm that AceTone's results are visually pleasing and stylistically coherent, demonstrating a new pathway toward language-driven, aesthetic-aligned color grading.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。