arXiv:2509.26006cs.CV2025-09被引 4

让图像质量评估像专家一样思考,自动分析并解释缺陷。

AgenticIQA: An Agentic Framework for Adaptive and Interpretable Image Quality Assessment

  • 用智能体框架拆解评估流程,动态选择工具
  • 在多个数据集上评分更准,解释更贴近人眼感知
  • 适合需要可解释性评估的科研与工业场景

图像质量评估(IQA)本质复杂,既需量化又需解释,源于人类视觉系统。传统方法依赖固定模型输出单一分数,难以适应多样失真、用户需求及解释要求。评分与解释常被割裂,而二者实则相辅相成:解释揭示感知退化,评分则将其浓缩为指标。为此,我们提出AgenticIQA,一种模块化智能体框架,将视觉语言模型(VLMs)与传统IQA工具动态结合,实现查询感知的评估。该框架将IQA分解为四个子任务:失真检测、分析、工具选择与执行,由规划器、执行器和总结器协同完成。规划器制定策略,执行器通过调用工具收集感知证据,总结器整合证据生成准确分数与人类对齐的解释。为支持训练与评估,我们构建了AgenticIQA-200K大规模指令数据集,以及首个用于评估VLM-based IQA智能体规划、执行与总结能力的基准AgenticIQA-Eval。跨多种IQA数据集的实验表明,AgenticIQA在评分准确性和解释一致性上均持续优于强基线。

原文摘要 · Abstract (English)

Image quality assessment (IQA) is inherently complex, as it reflects both the quantification and interpretation of perceptual quality rooted in the human visual system. Conventional approaches typically rely on fixed models to output scalar scores, limiting their adaptability to diverse distortions, user-specific queries, and interpretability needs. Furthermore, scoring and interpretation are often treated as independent processes, despite their interdependence: interpretation identifies perceptual degradations, while scoring abstracts them into a compact metric. To address these limitations, we propose AgenticIQA, a modular agentic framework that integrates vision-language models (VLMs) with traditional IQA tools in a dynamic, query-aware manner. AgenticIQA decomposes IQA into four subtasks -- distortion detection, distortion analysis, tool selection, and tool execution -- coordinated by a planner, executor, and summarizer. The planner formulates task-specific strategies, the executor collects perceptual evidence via tool invocation, and the summarizer integrates this evidence to produce accurate scores with human-aligned explanations. To support training and evaluation, we introduce AgenticIQA-200K, a large-scale instruction dataset tailored for IQA agents, and AgenticIQA-Eval, the first benchmark for assessing the planning, execution, and summarization capabilities of VLM-based IQA agents. Extensive experiments across diverse IQA datasets demonstrate that AgenticIQA consistently surpasses strong baselines in both scoring accuracy and explanatory alignment.

图像质量智能体可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。