arXiv:2606.30026cs.CVcs.AI2026-06被引 1

评测大模型对艺术创作意图的理解能力,发现当前最好模型仅达48%准确率。

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

论文配图:MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
图 1 · 摘自论文原文
  • 从1万份视频评论中提炼4016个问题,覆盖影视、视觉、表演、游戏四类艺术
  • 模型最高准确率48.29%,远低于人类专家的87.18%
  • 采用多轮迭代筛选+对抗性干扰项,提升艺术理解评估的严谨性

音频视觉艺术涵盖电影、视觉艺术、舞台表演和游戏设计等多种创作领域,艺术意义源于视觉、听觉与叙事元素的精心组合(如通过狭小空间构图强化恐惧感,或通过沉默与长时间特写传递悲伤)。真正的艺术理解不仅在于识别内容,更在于推理其创作意图。尽管多模态大语言模型(MLLMs)进展迅速,但现有基准主要衡量感知识别,忽视对创作意图的推理。为此,我们提出MuseBench,一个全面评估MLLM在艺术理解方面细微能力的基准。该基准包含4,016道题目,覆盖电影艺术、静态视觉艺术、舞台表演艺术和游戏艺术,源自超过10,000份配以专业解说与视觉示范的视频论文。为捕捉艺术分析的开放性,基准融合单选与可变选项多选题。所有题目通过四阶段迭代流程生成与优化,包括捷径过滤、对抗性干扰项设计及专家验证。对28个前沿MLLM的零样本评估显示,最先进模型准确率仅为48.29%,显著低于人类专家的87.18%,揭示当前模型在创意领域知识上的巨大差距。

原文摘要 · Abstract (English)

Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.g., fear amplified through claustrophobic framing, or grief conveyed through silence and lingering close-ups). True artistic understanding extends beyond recognizing what is depicted to reasoning about why it is expressed through particular creative choices. Despite the strong progress of multimodal large language models (MLLMs), this critical aspect of artistic understanding remains underexplored, as existing benchmarks largely measure perceptual recognition while overlooking reasoning about creative intent. To address this gap, we introduce Musebench, a comprehensive benchmark designed to evaluate MLLMs on nuanced artistic understanding. It comprises 4,016 questions spanning cinematic arts, static visual arts, stage performing arts, and game arts, distilled from over 10K candidate video essays that pair professional commentary with visual demonstration. To capture the open-ended nature of artistic analysis at scale, the benchmark combines single-select and variable-option multi-select questions. All questions are generated and refined through a four-phase iterative pipeline combining shortcut filtering, adversarial distractors, and expert validation. Comprehensive zero-shot evaluation of 28 state-of-the-art MLLMs reveals that even the best-performing model achieves only 48.29% accuracy, substantially below human expert performance of 87.18%, exposing a significant gap in current models' creative domain expertise.

艺术理解多模态模型评估基准创作意图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。