arXiv:2603.10990cs.CV2026-03中稿 · CVPR被引 1

评测并提升文本生成图像的色彩真实感,避免过度饱和失真。

Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity

  • 构建130万张真实与合成图像数据集,量化色彩真实度等级。
  • 提出多模态度量模型,自动评估生成图像色彩真实性。
  • 无需训练的调优方法,动态调节生成过程提升色彩可信度。

文本到图像(T2I)生成技术虽大幅提升视觉质量,但生成图像仍难达到真实摄影的自然感。现有评估方式(如人工评分和偏好训练指标)倾向于高饱和、高对比度的视觉效果,导致生成结果常因过度鲜艳而失真,即便用户提示追求写实风格。为此,我们构建了颜色保真度数据集(CFD),包含超过130万张具有有序色彩真实度等级的真实与合成图像,并提出颜色保真度度量(CFM),采用多模态编码器学习感知色彩保真度。同时,我们设计了一种无训练的颜色保真度优化(CFR)方法,通过自适应调节生成过程中的时空引导尺度,提升色彩真实性。CFD支持CFM评估,其学习到的注意力机制进一步指导CFR优化,形成一套渐进式评估与改进写实风格T2I生成色彩保真度的框架。代码与数据集已开源。

原文摘要 · Abstract (English)

Recent advances in text-to-image (T2I) generation have greatly improved visual quality, yet producing images that appear visually authentic to real-world photography remains challenging. This is partly due to biases in existing evaluation paradigms: human ratings and preference-trained metrics often favor visually vivid images with exaggerated saturation and contrast, which make generations often too vivid to be real even when prompted for realistic-style images. To address this issue, we present Color Fidelity Dataset (CFD) and Color Fidelity Metric (CFM) for objective evaluation of color fidelity in realistic-style generations. CFD contains over 1.3M real and synthetic images with ordered levels of color realism, while CFM employs a multimodal encoder to learn perceptual color fidelity. In addition, we propose a training-free Color Fidelity Refinement (CFR) that adaptively modulates spatial-temporal guidance scale in generation, thereby enhancing color authenticity. Together, CFD supports CFM for assessment, whose learned attention further guides CFR to refine T2I fidelity, forming a progressive framework for assessing and improving color fidelity in realistic-style T2I generation. The dataset and code are available at https://github.com/ZhengyaoFang/CFM.

图像生成色彩保真评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。