arXiv:2501.09012cs.CVcs.AI2025-01被引 20

让多模态大模型零样本判断艺术美感,提升与人类审美一致的推理能力。

Multimodal LLMs Can Reason about Aesthetics in Zero-Shot

  • 用证据驱动的思维链引导多模态大模型进行美学判断。
  • 新方法使模型推理更贴近人类审美,减少主观臆断和幻觉。
  • 适合艺术生成、AI美育和图像生成奖励建模场景。

生成艺术(GenArt)的技术飞速发展,使视觉上吸引人的图像创作变得普及。然而,真正具有艺术感染力——能引发观众深层共鸣的作品——仍难以实现,因其需要复杂的审美感知能力,而当前计算方法往往忽略这一多维度认知过程。本文首次探索如何有效激发多模态大模型(MLLMs)的推理能力以完成美学判断。分析发现,MLLMs在美学推理中易产生幻觉,表现为主观意见和无依据的艺术解读。我们进一步证明,通过采用基于证据和客观性的推理流程,可显著抑制此类幻觉。我们提出的基准方法 ArtCoT 促使模型生成多维度、深入的美学推理,与人类判断显著更一致。该成果可直接应用于AI艺术教学及图像生成的奖励建模。我们希望本工作为能够真正理解、欣赏并贡献于契合人类审美价值的艺术生成系统铺平道路。

原文摘要 · Abstract (English)

The rapid technical progress of generative art (GenArt) has democratized the creation of visually appealing imagery. However, achieving genuine artistic impact - the kind that resonates with viewers on a deeper, more meaningful level - remains formidable as it requires a sophisticated aesthetic sensibility. This sensibility involves a multifaceted cognitive process extending beyond mere visual appeal, which is often overlooked by current computational methods. This paper pioneers an approach to capture this complex process by investigating how the reasoning capabilities of Multimodal LLMs (MLLMs) can be effectively elicited to perform aesthetic judgment. Our analysis reveals a critical challenge: MLLMs exhibit a tendency towards hallucinations during aesthetic reasoning, characterized by subjective opinions and unsubstantiated artistic interpretations. We further demonstrate that these hallucinations can be suppressed by employing an evidence-based and objective reasoning process, as substantiated by our proposed baseline, ArtCoT. MLLMs prompted by this principle produce multifaceted, in-depth aesthetic reasoning that aligns significantly better with human judgment. These findings have direct applications in areas such as AI art tutoring and as reward models for image generation. Ultimately, we hope this work paves the way for AI systems that can truly understand, appreciate, and contribute to art that aligns with human aesthetic values. Project homepage: https://github.com/songrise/MLLM4Art.

多模态美学判断生成艺术推理链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。