arXiv:2602.20853cs.CV2026-02被引 2

用可解释AI分析视觉语言模型在艺术史中的推理逻辑。

On the Explainability of Vision-Language Models in Art History

  • 测试七种XAI方法,结合零样本定位与人工评估
  • 模型对概念稳定、表征清晰的类别更易解释
  • 适合关注艺术史与AI可解释性的研究者

视觉语言模型(VLMs)将视觉与文本数据映射到共享嵌入空间,支持多种多模态任务,也引发关于机器‘理解’本质的思考。本文研究可解释人工智能(XAI)方法如何揭示VLM——以CLIP为例——在艺术史语境下的视觉推理过程。通过结合零样本定位实验与人类可解释性研究,评估七种XAI方法的表现。结果表明,尽管这些方法能捕捉部分人类解读特征,其有效性依赖于所考察类别的概念稳定性与表征可用性。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) transfer visual and textual data into a shared embedding space. In so doing, they enable a wide range of multimodal tasks, while also raising critical questions about the nature of machine 'understanding.' In this paper, we examine how Explainable Artificial Intelligence (XAI) methods can render the visual reasoning of a VLM - namely, CLIP - legible in art-historical contexts. To this end, we evaluate seven methods, combining zero-shot localization experiments with human interpretability studies. Our results indicate that, while these methods capture some aspects of human interpretation, their effectiveness hinges on the conceptual stability and representational availability of the examined categories.

可解释AI艺术史视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。