让大模型同时看图、评质、解释,实现更懂人的图像质量评估。
iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
- 用统一多模态大模型同步完成质量定位、感知与描述。
- 在ViDA-UGC数据集上三项任务均达顶尖水平,获ICCV挑战赛第一。
- 适合需要可解释图像质量分析的科研与工业场景。
图像质量评估(IQA)已从单一数值预测发展为更具可解释性、贴近人类认知的新范式。本文针对细致化与可解释性IQA的新兴挑战,提出iDETEX——一种统一的多模态大语言模型(MLLM),可同步执行三大核心任务:质量定位、感知判断与详细描述。为实现跨异构子任务的高效泛化训练,设计了一套任务专用离线增强模块与数据混合策略,并辅以在线增强机制,充分挖掘多源监督信号。在大规模ViDA-UGC基准测试中,iDETEX在所有子任务上均达到领先性能,在ICCV MIPI 2025详细图像质量评估挑战赛中排名第一,验证了其在提供准确且可解释质量评估方面的有效性与鲁棒性。
原文摘要 · Abstract (English)
Image Quality Assessment (IQA) has progressed from scalar quality prediction to more interpretable, human-aligned evaluation paradigms. In this work, we address the emerging challenge of detailed and explainable IQA by proposing iDETEX-a unified multimodal large language model (MLLM) capable of simultaneously performing three key tasks: quality grounding, perception, and description. To facilitate efficient and generalizable training across these heterogeneous subtasks, we design a suite of task-specific offline augmentation modules and a data mixing strategy. These are further complemented by online enhancement strategies to fully exploit multi-sourced supervision. We validate our approach on the large-scale ViDA-UGC benchmark, where iDETEX achieves state-of-the-art performance across all subtasks. Our model ranks first in the ICCV MIPI 2025 Detailed Image Quality Assessment Challenge, demonstrating its effectiveness and robustness in delivering accurate and interpretable quality assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。