arXiv:2606.16082cs.CVcs.AI2026-06被引 1

让AI像人一样用放大镜和调光工具检查图片质量,效果更准。

Tool-IQA: Augmenting Image Quality Assessment with Simple Tools

论文配图:Tool-IQA: Augmenting Image Quality Assessment with Simple Tools
图 1 · 摘自论文原文
  • 给视觉语言模型加放大镜和调光工具,动态检查细节和隐藏瑕疵。
  • 在CLIVE数据集上相关系数达0.854,显著优于现有方法。
  • 适合需要高精度图像质量评估的研究与工业应用。

视觉语言模型(VLM)被越来越多地用于图像质量评估(IQA)。然而,当前方法通常采用静态的一次性评分范式,而人类评估图像质量时会通过动态观察,如选择性调整视角来验证细节和细微瑕疵。仅依赖单次观察带来两大局限:一是仅从全局尺度感知,难以评估局部细节;二是原始亮度分布可能掩盖可见性,导致对质量的检查不足。为此,我们提出Tool-IQA,将评估机制从被动打分转向工具增强的工作流。具体地,为VLM配备简单但有效的视图工具:放大镜用于检查局部细节,伽马校正器用于揭示可见性及隐藏瑕疵。评估流程包括初始观察并记录要点、工具辅助的深度检查、最终量化生成校准的质量分数。为进一步确保工具调用的高效性和目的性,我们引入批次感知训练策略,奖励能产生实际贡献的工具交互,而非单纯鼓励使用。在多个IQA基准测试中,实验表明,通过有效工具调用与校准评估,所提出的Tool-IQA显著优于现有最先进模型,例如在具有挑战性的CLIVE数据集上达到PLCC 0.854。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have been increasingly adopted for Image Quality Assessment (IQA). However, current methods typically employ a static one-shot scoring paradigm, despite the fact that humans assess image quality through dynamic visual inspection, e.g., selectively adjusting views to verify details and subtle artifacts. Specifically, relying solely on a single-pass observation introduces two primary limitations: first, perceiving the image only at a global scale restricts the assessment of finer local details; second, the original intensity distribution of the image may overwhelm the visibility, leading to insufficient inspection of image quality. To address these issues, we propose Tool-IQA, shifting the assessment mechanism from passive scoring to a tool-augmented workflow. In particular, we equip VLMs with simple yet effective view tools: a Magnifier to inspect local details, and a Gamma Corrector to uncover visibility and hidden artifacts. The assessment follows a structured pipeline that consists of an initial observation with rubric notes, a tool-augmented in-depth inspection, and a final quantification for calibrated quality score. Furthermore, to ensure efficient and purposeful tool callings, we introduce a batch-aware training strategy to reward tool interactions that can yield positive contributions rather than simply encouraging usage. Experiments on a variety of IQA benchmarks demonstrate that, with effective tool calling and calibrated assessment, our proposed Tool-IQA significantly outperforms existing state-of-the-art models, e.g., it achieves a PLCC of 0.854 on the challenging CLIVE dataset.

图像质量视觉语言模型工具增强评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。