arXiv:2608.01730cs.CV2026-08中稿 · the IEEE/CVF Confe…

新模型可跨分辨率评估图像质量,不降分辨率也不增加算力。

Learning Where to Look and How to Judge: Resolution-agnostic Image Quality Assessment with Quality-aware Saliency

论文配图:Learning Where to Look and How to Judge: Resolution-agnostic Image Quality Assessment with Quality-aware Saliency
图 1 · 摘自论文原文
  • 用多尺度补丁和注意力机制决定看哪里、怎么判断质量
  • 在多种图像数据集上超越现有方法,且计算量更低
  • 适合需要快速、通用图像质量评估的开发者

无参考图像质量评估(NR IQA)近期受益于深度多模态模型,但许多先进系统仍存在缺陷:或通过大幅缩放丢弃关键质量线索,或无法跨分辨率泛化,或难以在主观评分尺度不一的数据集上联合训练,或需极高算力。我们提出 ReLIQS——一种分辨率无关的图像质量学习模型,能保留原始分辨率的质量线索,从多个主观研究中学习,并保持计算高效与预算自适应。ReLIQS 是基于 CLIP 的多尺度补丁驱动架构,同时学习‘看哪里’和‘如何判断’质量。固定大小的补丁在多个分辨率(包括原始分辨率)上采样,经由 CLIP 视觉主干编码;轻量级感知重要性估计器生成特定于 IQA 的重要性图,筛选出少量高信息量补丁;隐空间质量轴模块将其嵌入聚合为单个图像级评分。在涵盖真实、合成及 AIGC 图像、多种分辨率与失真类型的基准测试中,ReLIQS 在性能上优于主流的 CNN、CLIP 及 MLLM 基线模型,且计算成本相当或更低。

原文摘要 · Abstract (English)

No-reference image quality assessment (NR IQA) has recently benefited from deep and multimodal models, yet many SOTA systems still violate at least one basic requirement: they either discard critical quality cues via aggressive resizing, fail to generalize across resolutions, cannot be jointly trained on heterogeneous IQA datasets with mismatched MOS scales, or require prohibitive computation. We present \textbf{ReLIQS}, a model for \textbf{Re}solution-agnostic \textbf{L}earning for \textbf{I}mage \textbf{Q}uality with \textbf{S}aliency, which is resolution-agnostic, preserves original-resolution quality cues, learns from multiple subjective studies, and remains computationally efficient and budget-adaptive. ReLIQS is a CLIP-based multiscale patch-driven architecture that learns both \emph{where to look} and \emph{how to judge} quality. Fixed-size patches are sampled across multiple resolutions, including the original resolution, and encoded with a CLIP vision backbone. A lightweight Perceptual Importance Estimator then predicts IQA-specific importance maps to select a small set of informative patches, and a Latent Quality Axis Module aggregates their embeddings into a single image-level score. Across authentic, synthetic, and AIGC benchmarks spanning diverse resolutions and distortions, ReLIQS generalizes better than strong CNN-, CLIP-, and MLLM-based baselines with matching or reduced computational cost.

图像质量评估多尺度分析CLIP应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。