评测大模型对图像质量的多层级判断能力,推动视觉质量评估向人类对齐。
VQualA 2025 Challenge on Visual Quality Comparison for Large Multimodal Models: Methods and Results
- 构建千级细粒度图像质量对比任务,涵盖单图、成对及多图场景。
- 采用2AFC和多选题等综合评估方式,验证模型在复杂质量判断中的表现。
- 5个指令微调模型展现新能力,适合关注可解释性评估的研究者。
本文总结了在ICCV 2025视觉质量评估研讨会期间举办的VQualA 2025挑战赛,旨在评估和提升主流大语言-视觉模型(LMMs)对多图像间视觉质量差异进行开放式、精细化推理的能力。挑战赛引入了一个新型基准,包含数千个从粗到细的视觉质量对比任务,覆盖单图、成对图像及多图像组。每项任务要求模型给出准确的质量判断。评估采用2AFC二选一偏好测试与多选题(MCQs)等综合协议。约100名参与者提交结果,其中五个模型展现出指令微调后在质量评估任务上的新兴能力。该挑战为开放域视觉质量推理提供了重要进展,并成为未来可解释性与人类对齐质量评价系统研究的催化剂。
原文摘要 · Abstract (English)
This paper presents a summary of the VQualA 2025 Challenge on Visual Quality Comparison for Large Multimodal Models (LMMs), hosted as part of the ICCV 2025 Workshop on Visual Quality Assessment. The challenge aims to evaluate and enhance the ability of state-of-the-art LMMs to perform open-ended and detailed reasoning about visual quality differences across multiple images. To this end, the competition introduces a novel benchmark comprising thousands of coarse-to-fine grained visual quality comparison tasks, spanning single images, pairs, and multi-image groups. Each task requires models to provide accurate quality judgments. The competition emphasizes holistic evaluation protocols, including 2AFC-based binary preference and multi-choice questions (MCQs). Around 100 participants submitted entries, with five models demonstrating the emerging capabilities of instruction-tuned LMMs on quality assessment. This challenge marks a significant step toward open-domain visual quality reasoning and comparison and serves as a catalyst for future research on interpretable and human-aligned quality evaluation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。