首个专用于人像图像质量感知的多模态大模型评测基准。
Q-Bench-Portrait: Benchmarking Multimodal Large Language Models on Portrait Image Quality Perception
- 构建涵盖2765组图文问答的综合性人像质量评估数据集。
- 覆盖技术失真、AI生成特有失真和美学评价等多维度指标。
- 适合研究人像生成与感知的学者及模型开发者参考。
多模态大语言模型在通用图像低级视觉任务上表现优异,但对具有独特结构与感知特性的肖像图像质量感知能力仍缺乏系统评估。为此,我们提出Q-Bench-Portrait,首个面向人像图像质量感知的综合性基准,包含2,765组图像-问题-答案三元组,涵盖自然、合成失真、AI生成、艺术及计算机图形等多种图像来源;评估维度包括技术失真、AIGC特有失真与美学;题型多样,支持单选、多选、判断与开放式问题,覆盖全局与局部层级。基于该基准,我们评估了20个开源与5个闭源多模态大模型,发现当前模型在人像感知方面虽具一定能力,但整体表现有限且不精确,与人类判断存在明显差距。本工作旨在推动通用与专用多模态模型在人像感知领域的进步。
原文摘要 · Abstract (English)
Recent advances in multimodal large language models (MLLMs) have demonstrated impressive performance on existing low-level vision benchmarks, which primarily focus on generic images. However, their capabilities to perceive and assess portrait images, a domain characterized by distinct structural and perceptual properties, remain largely underexplored. To this end, we introduce Q-Bench-Portrait, the first holistic benchmark specifically designed for portrait image quality perception, comprising 2,765 image-question-answer triplets and featuring (1) diverse portrait image sources, including natural, synthetic distortion, AI-generated, artistic, and computer graphics images; (2) comprehensive quality dimensions, covering technical distortions, AIGC-specific distortions, and aesthetics; and (3) a range of question formats, including single-choice, multiple-choice, true/false, and open-ended questions, at both global and local levels. Based on Q-Bench-Portrait, we evaluate 20 open-source and 5 closed-source MLLMs, revealing that although current models demonstrate some competence in portrait image perception, their performance remains limited and imprecise, with a clear gap relative to human judgments. We hope that the proposed benchmark will foster further research into enhancing the portrait image perception capabilities of both general-purpose and domain-specific MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。