用知识图谱增强艺术理解,让大模型更懂画作背后的文化背景。
ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
- 构建艺术知识图谱,整合艺术家、流派、主题等多维度信息。
- 在两个数据集上超越多个训练模型,生成更连贯深刻的解读。
- 无需训练,适合需要文化语境理解的艺术分析场景。
理解视觉艺术需跨越文化、历史与风格等多重视角,远超物体识别范畴。尽管当前多模态大模型在通用图像描述上表现良好,却难以捕捉精细艺术所需的深层诠释。我们提出ArtRAG,一种无需训练的框架,结合结构化知识与检索增强生成(RAG),实现多视角艺术解释。ArtRAG自动从领域文本中构建艺术上下文知识图谱(ACKG),将艺术家、艺术流派、主题与历史事件等实体组织为可解释的复杂网络。推理时,多粒度结构化检索器选取语义与拓扑相关的子图以引导生成,使多模态大模型产出具有情境依据与文化深度的艺术描述。在SemArt与Artpedia数据集上的实验表明,ArtRAG优于多个经大量训练的基线模型;人工评估进一步证实其生成内容连贯、富有洞见且文化内涵丰富。
原文摘要 · Abstract (English)
Understanding visual art requires reasoning across multiple perspectives -- cultural, historical, and stylistic -- beyond mere object recognition. While recent multimodal large language models (MLLMs) perform well on general image captioning, they often fail to capture the nuanced interpretations that fine art demands. We propose ArtRAG, a novel, training-free framework that combines structured knowledge with retrieval-augmented generation (RAG) for multi-perspective artwork explanation. ArtRAG automatically constructs an Art Context Knowledge Graph (ACKG) from domain-specific textual sources, organizing entities such as artists, movements, themes, and historical events into a rich, interpretable graph. At inference time, a multi-granular structured retriever selects semantically and topologically relevant subgraphs to guide generation. This enables MLLMs to produce contextually grounded, culturally informed art descriptions. Experiments on the SemArt and Artpedia datasets show that ArtRAG outperforms several heavily trained baselines. Human evaluations further confirm that ArtRAG generates coherent, insightful, and culturally enriched interpretations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。