arXiv:2605.09479eess.IVcs.CV2026-05

用多模型一致性评估图像质量,更符合机器任务需求。

ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality

论文配图:ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality
图 1 · 摘自论文原文
  • 基于多个预训练模型的一致性投票构建数据集
  • 通过融合局部与全局特征提升机器感知质量判断能力
  • 适合图像压缩、下游任务优化等机器学习场景

我们从机器视角研究全参考图像质量评估,将机器导向的质量定义为潜在的机器效用,并通过成对预测一致性比较进行近似。为此,我们构建了PCMP数据集——一组保真度匹配的失真图像对,由多个预训练模型的一致性投票标注。我们进一步提出ML-CLIPSim,一种基于冻结的CLIP视觉编码器的可微质量度量,该方法聚合中间块-标记相似性与全局图像嵌入。在机器偏好基准、人类图像质量评估数据集以及学习型图像压缩上的实验表明,ML-CLIPSim相较于传统保真度和感知指标更契合机器导向偏好,同时保持对人类质量预测的竞争力。作为压缩失真项使用时,它在多个下游任务中提升了率-任务权衡表现。

原文摘要 · Abstract (English)

We study full-reference image quality assessment from a machine-centric perspective, where images are evaluated by how well they preserve information for downstream models. We formulate machine-oriented quality as a latent machine utility and approximate it through pairwise predictive-consistency comparisons. To this end, we construct PCMP, a dataset of PSNR-matched distortion pairs labeled by consistency votes from multiple pretrained models. We further propose ML-CLIPSim, a differentiable quality metric built on a frozen CLIP visual encoder, which aggregates intermediate patch-token similarities and global image embeddings. Experiments on machine-preference benchmarks, human-IQA datasets, and learned image compression show that ML-CLIPSim better aligns with machine-oriented preferences than conventional fidelity and perceptual metrics, while remaining competitive for human quality prediction. Used as a compression distortion term, it improves rate--task trade-offs across multiple downstream tasks.

图像质量评估机器感知视觉编码器压缩优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。