用单个词预测数值,让大模型轻松评估图片质量与美感
Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model
- 仅预测两个额外数字即可实现最优评分效果
- 在5个公开数据集上超越现有方法,零样本泛化能力强
- 构建1.47万张UGC图片的细粒度标注数据集
移动互联网的快速发展导致用户生成内容(UGC)图像大幅增加,对UGC图像的全面评估变得迫切且重要。近年来,多模态大语言模型(MLLMs)在图像质量评估(IQA)和图像美学评估(IAA)方面展现出巨大潜力。然而,有效评估UGC图像的质量与美感仍面临两大挑战:1)单一分数无法捕捉人类感知的层次性;2)如何利用MLLM输出数值评分(如平均意见分,MOS)仍是开放问题。为此,我们提出一个新数据集RealQA,包含14,715张UGC图像,每张图像均标注了10个细粒度属性,涵盖低层(如图像清晰度)、中层(如主体完整性)和高层(如构图)三个层级。此外,我们深入研究了如何有效利用MLLM预测数值评分。令人惊讶的是,通过仅预测两个额外显著数字,即“下一个词”范式,即可达到当前最佳性能。结合思维链(CoT)与学习到的细粒度属性,所提方法在五个公开IQA与IAA数据集上表现优于现有方法,具备更强可解释性,并在视频质量评估(VQA)上展现优异零样本泛化能力。代码与数据集将公开。
原文摘要 · Abstract (English)
The rapid expansion of mobile internet has resulted in a substantial increase in user-generated content (UGC) images, thereby making the thorough assessment of UGC images both urgent and essential. Recently, multimodal large language models (MLLMs) have shown great potential in image quality assessment (IQA) and image aesthetic assessment (IAA). Despite this progress, effectively scoring the quality and aesthetics of UGC images still faces two main challenges: 1) A single score is inadequate to capture the hierarchical human perception. 2) How to use MLLMs to output numerical scores, such as mean opinion scores (MOS), remains an open question. To address these challenges, we introduce a novel dataset, named Realistic image Quality and Aesthetic (RealQA), including 14,715 UGC images, each of which is annoted with 10 fine-grained attributes. These attributes span three levels: low level (e.g., image clarity), middle level (e.g., subject integrity) and high level (e.g., composition). Besides, we conduct a series of in-depth and comprehensive investigations into how to effectively predict numerical scores using MLLMs. Surprisingly, by predicting just two extra significant digits, the next token paradigm can achieve SOTA performance. Furthermore, with the help of chain of thought (CoT) combined with the learnt fine-grained attributes, the proposed method can outperform SOTA methods on five public datasets for IQA and IAA with superior interpretability and show strong zero-shot generalization for video quality assessment (VQA). The code and dataset will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。