arXiv:2601.15356eess.IVcs.AI2026-01

让AI看高清图更懂细节,避免误判自然模糊为缺陷。

Q-Probe: Scaling Image Quality Assessment to High Resolution via Context-Aware Agentic Probing

  • 用智能探针式分析局部图像,结合上下文判断质量
  • 在高分辨率下超越现有模型,且跨尺度表现稳定
  • 适合需要精准图像质量评估的研究与工业场景

强化学习使多模态大模型在图像质量评估中逼近人类偏好,但现有方法依赖粗粒度全局视图,难以捕捉高分辨率下的细微局部退化。尽管新兴的“以图思考”范式通过缩放机制实现多尺度视觉感知,但直接应用于图像质量评估会引入‘裁剪即退化’的错误偏见,并将自然景深误判为伪影。为此,我们提出 Q-Probe,首个基于上下文感知探针的高分辨率图像质量评估框架。首先构建 Vista-Bench,这是首个面向高分辨率图像质量评估中细粒度局部退化分析的基准。其次,提出三阶段训练范式,在逐步对齐人类偏好的同时,通过新型上下文感知裁剪策略消除因果偏差。大量实验表明,Q-Probe 在高分辨率设置下达到领先性能,并在多分辨率尺度上保持优异效果。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has empowered Multimodal Large Language Models (MLLMs) to achieve superior human preference alignment in Image Quality Assessment (IQA). However, existing RL-based IQA models typically rely on coarse-grained global views, failing to capture subtle local degradations in high-resolution scenarios. While emerging "Thinking with Images" paradigms enable multi-scale visual perception via zoom-in mechanisms, their direct adaptation to IQA induces spurious "cropping-implies-degradation" biases and misinterprets natural depth-of-field as artifacts. To address these challenges, we propose Q-Probe, the first agentic IQA framework designed to scale IQA to high resolution via context-aware probing. First, we construct Vista-Bench, a pioneering benchmark tailored for fine-grained local degradation analysis in high-resolution IQA settings. Furthermore, we propose a three-stage training paradigm that progressively aligns the model with human preferences, while simultaneously eliminating causal bias through a novel context-aware cropping strategy. Extensive experiments demonstrate that Q-Probe achieves state-of-the-art performance in high-resolution settings while maintaining superior efficacy across resolution scales.

图像质量评估多模态模型强化学习高分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。