arXiv:2509.12750cs.CV2025-09AAAI

对比人类与多模态大模型对图像质量的判断差异。

What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference Alignment

  • 构建合成图像对数据集,分析人类与模型对图像属性的评价关联。
  • 发现人类对图像美学、构图等属性判断更一致,模型则关联较弱。
  • 模型在解剖准确度等细节上判断能力明显弱于人类,适合研究评估机制。

生成式文本到图像模型的自动化评估仍具挑战性。尽管近期研究提出使用多模态大模型评判图像质量,但对其如何利用人类关注的图像风格、构图等概念进行整体评估缺乏深入理解。本文研究了美学、无瑕疵、解剖准确性、构图正确性、对象一致性及风格等图像属性对人类与多模态大模型判断质量的重要性。首先,我们通过合成图像对构建人类偏好数据集,并分析各质量属性之间的任务相关性,发现人类判断中这些属性存在较强关联。重复相同分析发现,多模态大模型中的属性关联性显著更弱。进一步通过高控制力的合成数据集逐项考察各属性,结果表明人类能轻松识别各属性高低(如高/低美学图像),而部分属性如解剖准确度,对多模态大模型而言学习难度极大。综合来看,揭示了人类与多模态大模型在图像感知上的本质差异。

原文摘要 · Abstract (English)

Automated evaluation of generative text-to-image models remains a challenging problem. Recent works have proposed using multimodal LLMs to judge the quality of images, but these works offer little insight into how multimodal LLMs make use of concepts relevant to humans, such as image style or composition, to generate their overall assessment. In this work, we study what attributes of an image--specifically aesthetics, lack of artifacts, anatomical accuracy, compositional correctness, object adherence, and style--are important for both LLMs and humans to make judgments on image quality. We first curate a dataset of human preferences using synthetically generated image pairs. We use inter-task correlation between each pair of image quality attributes to understand which attributes are related in making human judgments. Repeating the same analysis with LLMs, we find that the relationships between image quality attributes are much weaker. Finally, we study individual image quality attributes by generating synthetic datasets with a high degree of control for each axis. Humans are able to easily judge the quality of an image with respect to all of the specific image quality attributes (e.g. high vs. low aesthetic image), however we find that some attributes, such as anatomical accuracy, are much more difficult for multimodal LLMs to learn to judge. Taken together, these findings reveal interesting differences between how humans and multimodal LLMs perceive images.

图像评估多模态人类对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。