评测并提升文本生成图像的色彩真实感,避免过度饱和失真。
Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity
- 构建130万张真实与合成图像数据集,量化色彩真实度等级。
- 提出多模态度量模型,自动评估生成图像色彩真实性。
- 无需训练的调优方法,动态调节生成过程提升色彩可信度。
文本到图像(T2I)生成技术虽大幅提升视觉质量,但生成图像仍难达到真实摄影的自然感。现有评估方式(如人工评分和偏好训练指标)倾向于高饱和、高对比度的视觉效果,导致生成结果常因过度鲜艳而失真,即便用户提示追求写实风格。为此,我们构建了颜色保真度数据集(CFD),包含超过130万张具有有序色彩真实度等级的真实与合成图像,并提出颜色保真度度量(CFM),采用多模态编码器学习感知色彩保真度。同时,我们设计了一种无训练的颜色保真度优化(CFR)方法,通过自适应调节生成过程中的时空引导尺度,提升色彩真实性。CFD支持CFM评估,其学习到的注意力机制进一步指导CFR优化,形成一套渐进式评估与改进写实风格T2I生成色彩保真度的框架。代码与数据集已开源。
原文摘要 · Abstract (English)
Recent advances in text-to-image (T2I) generation have greatly improved visual quality, yet producing images that appear visually authentic to real-world photography remains challenging. This is partly due to biases in existing evaluation paradigms: human ratings and preference-trained metrics often favor visually vivid images with exaggerated saturation and contrast, which make generations often too vivid to be real even when prompted for realistic-style images. To address this issue, we present Color Fidelity Dataset (CFD) and Color Fidelity Metric (CFM) for objective evaluation of color fidelity in realistic-style generations. CFD contains over 1.3M real and synthetic images with ordered levels of color realism, while CFM employs a multimodal encoder to learn perceptual color fidelity. In addition, we propose a training-free Color Fidelity Refinement (CFR) that adaptively modulates spatial-temporal guidance scale in generation, thereby enhancing color authenticity. Together, CFD supports CFM for assessment, whose learned attention further guides CFR to refine T2I fidelity, forming a progressive framework for assessing and improving color fidelity in realistic-style T2I generation. The dataset and code are available at https://github.com/ZhengyaoFang/CFM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。