评测大模型能否像人一样进行开放式审美批评,发现其表现与人类差异显著。
Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models

- 用8种提示条件对比8个开源多模态大模型与GPT-5.5的审美评语
- 模型评语与人类评语相似度低,主要受长度和风格影响而非图像内容
- 模型缺乏人类的审慎选择性,反而重复冗长,适合关注评语风格的研究者
开放式审美批评对多模态大语言模型(MLLMs)构成挑战:它没有单一正确答案,且多数评估方法依赖数值评分而非真实人类书面评语。本文评估了从70亿到3970亿参数的8个开源MLLM及GPT-5.5,在1,227条r/photocritique帖子上,基于八种提示条件生成的评语与人类评语的相似性。参考文本相似性带来误导;绝对指标中,词汇与学习型度量与人类评语关联微弱,仅粗粒度嵌入余弦值显示主题重叠。要求更短评语可提升分数,而移除图像几乎无影响:相似性反映的是长度、文本内容与稳定写作风格,而非图像特异性观察。使用LLM裁判进一步揭示矛盾:在主条件下,四个裁判均偏好前沿模型评语胜过人类,但在7–80亿参数模型上分歧巨大(9%至81%偏好),说明评分体系无法捕捉真实差异。当被问及实质相似性时,裁判与两名人类标注员一致认为所有模型得分在1.81至2.59间(1–5分制),接近“大多不同”。行为层面,模型覆盖几乎所有美学维度,而人类具有选择性;且同一照片下模型反复重复,即便按人类长度提示亦然。研究认为,参考式相似性更奖励流畅全面的风格,而非人类评语中的选择性与具体性。
原文摘要 · Abstract (English)
Open-ended aesthetic critique is a challenge for multimodal large language models (MLLMs): it has no single correct answer, and most aesthetic evaluation measures models against numeric scores rather than the written critiques people actually give. We ask whether MLLM critiques are close to human ones, scoring eight open-weight MLLMs from $7$B to $397$B, plus GPT-5.5, against multiple ranked human critiques for each of $1{,}227$ \texttt{r/photocritique} posts under eight prompt conditions. Reference-based similarity gives a misleading picture. In absolute terms the stricter lexical and learned metrics align only weakly with human critiques while a coarse embedding cosine reports broad topical overlap, yet requesting shorter critiques raises those scores and withholding the image barely changes them: the similarity reflects length, the post text, and a stable critiquing style more than image-specific observation. An LLM judge sharpens the question rather than settling it: in the primary condition all four judges prefer the frontier models' critiques to the human ones, but on the $7$--$8$B models they diverge wildly, from $9\%$ to $81\%$ preference on identical pairs. Asked instead how similar each pair is in substance, those judges and two human annotators agree, rating every model between $1.81$ and $2.59$ on a $1$--$5$ scale, close to ``mostly different''. Behaviorally, the models diverge in ways the scores do not surface: they cover nearly every aesthetic aspect where humans are selective and repeat themselves across critiques of one photo, even when prompted to write at human length. We argue that reference-based similarity rewards a fluent, comprehensive critique style rather than the selectivity and specificity of human critique.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。