arXiv:2606.29672cs.CL2026-06

无需训练,大模型可零样本评估视觉创意,且推理过程透明可解释。

How LLMs See Creativity: Zero-Shot Scoring of Visual Creativity with Interpretable Reasoning

论文配图:How LLMs See Creativity: Zero-Shot Scoring of Visual Creativity with Interpretable Reasoning
图 1 · 摘自论文原文
  • 用多模态大模型直接判断图像创意,不需微调或示例。
  • 在992张生成图和1500幅手绘图上,与人工评分相关性达0.57-0.68。
  • 模型推理过程清晰展示其关注点与评价逻辑,适合可信评估场景。

评估视觉图像的原创性长期面临挑战。尽管在语言领域自动化评分已有效,但对视觉创意的评估及模型评分机制的理解仍存疑问。本研究探究多模态大语言模型(LLMs)能否零样本(无需微调或人类评分示例)充当视觉创意裁判,并分析其推理输出是否提供可解释的评价路径。我们测试了六款多模态大模型(Gemini 3 Flash、Gemma 4 31B IT、GPT-5.4 Mini、GLM-5v Turbo、Kimi K2.5、Qwen 3.6 Plus)在992张基于人工提示生成的AI图像和1,500幅由人类评分的亲笔手绘图上的表现。研究1中,所有模型在两组数据上均与人类评分有显著相关性(生成图:r = .57–.68;手绘图:r = .29–.68)。研究2中,对三款模型的逐步推理过程分析显示,其推理使评价过程可解释——揭示了模型关注的特征、原创性与质量的权衡方式及评分依据,但推理本身未提升与人类评分的一致性。结果表明,多模态大模型可在无额外训练下匹配人类对视觉创意的判断,且其推理过程提供了理解模型评价机制的窗口。开放评分工具已在 https://review-visual-eval-scoring.hf.space 发布。

原文摘要 · Abstract (English)

Evaluating the originality of visual images poses enduring challenges for creativity assessment. Automated scoring using AI models has proven effective in the verbal domain, yet key questions remain about evaluating visual creativity and understanding how models arrive at their ratings. The present research asks whether multimodal large language models (LLMs) can serve as judges of visual creativity zero-shot (without any fine-tuning or examples of human ratings) and whether their "reasoning" output offers an interpretable window into their evaluation process. We tested six multimodal LLMs (Gemini 3 Flash, Gemma 4 31B IT, GPT-5.4 Mini, GLM-5v Turbo, Kimi K2.5, and Qwen 3.6 Plus) on 992 AI-generated images (based on human-written prompts) and 1,500 hand-drawn sketches scored for creativity by human raters. In Study 1, all models showed substantial alignment with human creativity ratings on both datasets (r = .57-.68 on AI-generated images; r = .29-68 on sketches). In Study 2, we analyzed the step-by-step reasoning processes of three LLMs evaluating the same images and drawings. Although reasoning made model evaluations interpretable -- showing what they attend to, how they balance originality vs. quality, and how they justify their ratings -- reasoning did not improve alignment with human ratings. In sum, our findings indicate that multimodal LLMs can match human judgments of visual creativity without any additional training, and that their reasoning reveals how AI models evaluate creativity. An open scoring app implementing this pipeline is available at https://review-visual-eval-scoring.hf.space.

视觉创意多模态模型零样本评估可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。