用记忆重排提升图像质量评估的敏感度,解决评分僵化问题。
ME-IQA: Memory-Enhanced Image Quality Assessment via Re-Ranking
- 构建记忆库,通过推理摘要检索语义感知邻居
- 融合成对偏好概率,使评分更密集且对失真敏感
- 动态更新记忆,适合高精度图像评估场景
基于文本推理的视觉语言模型在图像质量评估中表现优异,但其标量评分常因敏感性不足而陷入离散塌陷。本文提出ME-IQA,一种即插即用的测试时记忆增强重排框架。该方法(i)构建记忆库,利用推理摘要检索语义与感知对齐的邻居;(ii)将视觉语言模型重构为概率比较器,获取成对偏好概率,并在Thurstone's Case V模型下融合初始分数;(iii)执行门控反思并整合记忆以优化后续决策。该方法生成更密集、对失真更敏感的预测结果,有效缓解离散塌陷。在多个IQA基准上,相比强推理型VLM基线、传统非推理方法及测试时扩展方案均取得一致性能提升。
原文摘要 · Abstract (English)
Reasoning-induced vision-language models (VLMs) advance image quality assessment (IQA) with textual reasoning, yet their scalar scores often lack sensitivity and collapse to a few values, so-called discrete collapse. We introduce ME-IQA, a plug-and-play, test-time memory-enhanced re-ranking framework. It (i) builds a memory bank and retrieves semantically and perceptually aligned neighbors using reasoning summaries, (ii) reframes the VLM as a probabilistic comparator to obtain pairwise preference probabilities and fuse this ordinal evidence with the initial score under Thurstone's Case V model, and (iii) performs gated reflection and consolidates memory to improve future decisions. This yields denser, distortion-sensitive predictions and mitigates discrete collapse. Experiments across multiple IQA benchmarks show consistent gains over strong reasoning-induced VLM baselines, existing non-reasoning IQA methods, and test-time scaling alternatives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。