arXiv:2608.29644cs.CVcs.AI2026-08

用AI把画作视觉特征转为艺术史描述,让机器懂风格

Conducting Stylistic Analysis of Paintings through an Art-History Agent

  • 用ViT+字典学习提取跨作品共现风格特征
  • 通过LLM将特征关联画作与策展文本生成描述
  • 适合艺术史研究者和计算人文领域探索者

艺术品作者归属传统依赖细致的视觉观察与描述,即艺术史中的风格分析。当前人工智能模型在该领域仅提供无解释的概率分类。为弥合这一方法鸿沟,我们提出一个自动化风格分析框架,旨在增强证据收集、发现与验证。通过在含元数据的大规模绘画语料上训练视觉变换器(ViT),系统将艺术史特定数据编码为嵌入表示,并通过稀疏字典学习将其分解为一组共享特征,这些特征在训练集中反复出现。随后,大型语言模型(LLM)通过检索相关作品及其策展人撰写的文本,解读每个特征并合成反映其风格属性的描述。最后,一个自主协调的LLM采用推理-行动(ReAct)框架,对特征进行权重分配、测试与优化,形成对单幅作品的连贯描述或多作品比较。该方法将详细视觉特征转化为可解释的描述性术语,解决了艺术史中的关键挑战,实现了图像作为数据与人文学科语义关切的连接,确立了基于视觉的计算艺术史这一未来发展方向。

原文摘要 · Abstract (English)

Attributing an artwork to an artist has traditionally relied on detailed visual observations and descriptions, known as stylistic analysis in art history. By contrast, current artificial intelligence (AI) models used in the field offer only unexplained probabilistic classifications. To bridge this methodological gap, we present an AI framework that automates stylistic analysis of paintings, providing a foundation for enhancing evidence collection, discovery, and verification. By training a vision transformer (ViT) on a large corpus of paintings with metadata, our system encodes this art history-specific data as embeddings. These representations are factorized via sparse dictionary learning into a shared set of features that recur across the training set. A large language model (LLM) then interprets each feature by retrieving associated artworks and their accompanying curator-written texts, and synthesizes them into descriptions that reflect their stylistic attributes. Finally, an autonomous coordinator LLM applies a reasoning-and-action (ReAct) framework to weight, test, and refine these features into cohesive descriptions of an artwork, or comparisons of artworks. This approach converts detailed visual features into descriptive terms, addressing a key challenge in art history. It thus connects the use of images as data with the semantic concerns of humanists, establishing vision-based computational art history as an area for future growth.

艺术史风格分析多模态LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。