测试大模型在艺术评论与心理推理中的表现,发现其输出可媲美人类。
Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation
- 结合艺术理论构建提示框架,引导模型生成连贯深度的艺术评论。
- 人类难以区分AI与真人评论,证明模型可产出风格与内涵兼具的文本。
- 新设计的心理推理任务揭示模型在情感与道德情境下的认知差异,适合研究者参考。
本研究探讨大型语言模型(LLMs)在艺术领域的两项能力:艺术作品评论生成与艺术情境中的心智理论(ToM)推理。针对评论生成,我们融合诺埃尔·卡罗尔的评价框架与多元艺术批评理论,采用分步提示策略引导模型先生成完整评论,再提炼出更简洁连贯的版本。通过类图灵测试评估,多数人类评判者无法准确识别哪段评论由AI生成,表明在精心引导下,模型能产出风格逼真、解释丰富的评论。第二部分引入基于艺术情境的新型简单ToM任务,涵盖诠释、情感与道德张力,超越传统假信念测试,体现更复杂的社会化推理。测试了41个近期的LLMs,发现其表现因任务和模型而异,尤其在涉及情感或模糊情境的任务中差异更明显。整体结果揭示了大模型应对复杂阐释挑战时的认知局限与潜力,虽未直接否定生成式AI悖论(即模型可生成专家级输出却无真实理解),但表明通过精细提示设计,模型行为可能更接近理解。
原文摘要 · Abstract (English)
This study explored how large language models (LLMs) perform in two areas related to art: writing critiques of artworks and reasoning about mental states (Theory of Mind, or ToM) in art-related situations. For the critique generation part, we built a system that combines Noel Carroll's evaluative framework with a broad selection of art criticism theories. The model was prompted to first write a full-length critique and then shorter, more coherent versions using a step-by-step prompting process. These AI-generated critiques were then compared with those written by human experts in a Turing test-style evaluation. In many cases, human subjects had difficulty telling which was which, and the results suggest that LLMs can produce critiques that are not only plausible in style but also rich in interpretation, as long as they are carefully guided. In the second part, we introduced new simple ToM tasks based on situations involving interpretation, emotion, and moral tension, which can appear in the context of art. These go beyond standard false-belief tests and allow for more complex, socially embedded forms of reasoning. We tested 41 recent LLMs and found that their performance varied across tasks and models. In particular, tasks that involved affective or ambiguous situations tended to reveal clearer differences. Taken together, these results help clarify how LLMs respond to complex interpretative challenges, revealing both their cognitive limitations and potential. While our findings do not directly contradict the so-called Generative AI Paradox--the idea that LLMs can produce expert-like output without genuine understanding--they suggest that, depending on how LLMs are instructed, such as through carefully designed prompts, these models may begin to show behaviors that resemble understanding more closely than we might assume.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。