arXiv:2606.27947cs.CV2026-06中稿 · PRESTIGE workshop …

用激活图解析大模型如何看懂艺术画作,发现不同词汇依赖的视觉证据不同。

Understanding How MLLMs Describe Artworks Using Token Activation Maps

论文配图:Understanding How MLLMs Describe Artworks Using Token Activation Maps
图 1 · 摘自论文原文
  • 通过令牌激活图分离每个词对应的视觉证据,排除上下文干扰。
  • 模型识别艺术家准确率高于标题预测,后者更易出现幻觉。
  • 揭示风格、符号等不同语义词的视觉依赖模式差异,适合研究视觉推理者。

多模态大语言模型(MLLMs)能流畅描述艺术作品,但其输出背后的视觉推理机制仍不清晰。当模型命名风格、识别主题或辨认象征符号时,是基于画面特定区域,还是依赖整体视觉信号或文本先验?我们使用令牌激活图(TAM)分析这一问题,该方法为每个生成词生成热力图,仅保留与该词相关的特定视觉证据,排除上下文干扰。对涵盖多个时期和流派的精选画作集进行分析,考察五类语义不同的词:常见视觉对象、风格描述词、元数据、图像象征词和情感表达词的视觉锚定模式。结果表明,视觉锚定程度随词义变化显著。模型在艺术家归属任务中表现优于标题预测,后者幻觉更频繁。最后将TAM与SAM~3开放词汇分割对比。为确保可复现性,已公开代码、配置、提示词及结果。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) describe artworks with remarkable fluency, yet the visual reasoning behind their outputs remains opaque. When an MLLM names a style, identifies a subject, or recognizes an iconographic symbol, does it ground each claim in the relevant region of the canvas, draw on an undifferentiated visual signal, or rely primarily on textual priors? We study this using the Token Activation Map (TAM), which produces, for each generated token, a heatmap isolating the visual evidence specific to that token from prior-context interference. Applying TAM to a curated set of paintings spanning multiple periods and genres, we analyze grounding patterns across five semantically distinct token categories: common visual objects, style descriptors, metadata, iconographic tokens, and affective expressions. We find that visual grounding varies substantially with token semantics. We further show that MLLMs attempt to identify artworks and artists, achieving higher accuracy in artist attribution than in title prediction, where hallucinations are more frequent. Finally, we compare TAM with SAM~3 open-vocabulary segmentation. To ensure reproducibility, we release our code, experimental configurations, prompts, and qualitative results on the project page at https://nicolafan.github.io/tamart/.

多模态视觉推理艺术理解激活图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。