arXiv:2606.08833cs.CV2026-06

让图像生成更符合人眼感知,提升真实感与细节质量

CSFlow: Aligning Flow Matching with Human Contrast Sensitivity

论文配图:CSFlow: Aligning Flow Matching with Human Contrast Sensitivity
图 1 · 摘自论文原文
  • 基于人眼对比敏感度设计分频权重,引导生成顺序
  • 改进后FID降低4.7%,生成质量评分提升2.5%
  • 适合关注视觉真实性和生成细节的开发者

我们提出对比敏感流(CSFlow),一种将人类视觉系统的对比敏感度函数(CSF)与流匹配的迭代去噪过程相连接的加权策略。由于真实图像信号集中在低频成分,这些部分在连续扩散过程中较早达到高信噪比,导致傅里叶空间中出现软自回归结构:粗略内容先稳定,细节后生成。而人眼对不同空间频率的敏感度不均:极低和极高频率需更高对比度才能感知。我们首次通过两项贡献整合该现象:(1) 估计每个逆向流阶段生成的频率;(2) 根据各噪声水平下生成频率与人眼敏感度的匹配程度,获取时间步权重。实验验证表明,仅通过推理时调整时间步权重或短时微调,即可使FID降低4.7%,Inception Score提升2.2%,GenEval得分提高2.5%。定性上,采用CSFlow生成的图像更具视觉真实感,减少卡通化倾向。

原文摘要 · Abstract (English)

We introduce Contrast Sensitive Flow (CSFlow), a weighting scheme that connects the human eye's Contrast Sensitivity Function (CSF) to the iterative denoising steps of flow matching. Because real-world images concentrate signal at low spatial frequencies, these components reach high signal-to-noise ratio earlier during continuous diffusion than high-frequency components. When generating images with diffusion or flow matching models, this induces a soft autoregressive structure in Fourier space, where coarse image content stabilizes before fine detail. Meanwhile, the human visual system is unequally sensitive to spatial frequencies: very low and very high frequencies require significantly higher contrast to be perceived. We for the first time merge these observations through two contributions: (1) a metric that estimates which frequencies are generated at each reverse flow interval and (2) timestep weights obtained by aligning the frequencies generated at each noise level with human contrast sensitivity. We validate our contributions experimentally showing that these weights can improve generative performance by lowering FID by 4.7%, increasing Inception Score by 2.2% and improving GenEval scores by 2.5% using inference-only timestep modification or short fine-tuning. Qualitatively, we find that our CSFlow weights lead to better visual realism and less cartoonish appearance of generated images.

图像生成视觉感知流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。