让AI看图评质量,用工具生成证据提升判断可信度
IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment

- 用专用工具提取噪声、梯度等视觉证据辅助判断
- 在7个数据集上表现最优,且结果可解释
- 适合需要透明化质量评估的场景
开放世界中的图像质量评估仍面临泛化能力弱和可解释性差的挑战。基于多模态大语言模型(MLLM)的方法虽引入文本推理,但其判断依赖语义偏倚的内部表征,对低层感知退化不敏感。我们提出IQA-T1,一种基于工具的视觉证据推理框架,通过显式感知观察增强MLLM推理。推理时,模型自主调用专用分析工具生成结构化视觉证据,如噪声残差图、梯度统计量和频谱分布,并逐步融入推理过程。为支持该范式,我们构建了Q-Tool数据集,包含11,000条基于工具生成证据的多模态推理链。在七个IQA基准上的实验表明,IQA-T1在跨数据集表现最佳,且生成的评估结果具可解释性和证据支撑。代码与数据集已公开于https://github.com/zibuyu-02/IQA-T1。
原文摘要 · Abstract (English)
Image Quality Assessment (IQA) in open-world environments remains challenging due to limited generalization and interpretability. Recent approaches based on multimodal large language models (MLLMs) introduce textual reasoning for quality prediction, yet their judgments rely heavily on semantically biased internal representations, making them insensitive to low-level perceptual degradations. We propose IQA-T1, a tool-based visual evidence reasoning framework that augments MLLM reasoning with explicit perceptual observations. During inference, the model autonomously invokes specialized analysis tools to generate structured visual evidence, such as noise residual maps, gradient statistics, and frequency spectra, which are progressively integrated into the reasoning process. To support this paradigm, we construct Q-Tool, a dataset containing 11k multimodal reasoning chains grounded in tool-generated evidence. Extensive experiments on seven IQA benchmarks show that IQA-T1 achieves the best overall performance across datasets while producing interpretable and evidence-grounded quality assessments. Code and dataset are available at https://github.com/zibuyu-02/IQA-T1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。