用艺术元数据增强图像描述,让AI更懂画作深层含义。
KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph
- 融合元数据与异构知识图谱,提升艺术图像理解
- 在多个数据集上CIDEr得分超越现有最佳模型
- 适合艺术分析、数字策展等需要深度解读的场景
探索精美绘画所传达的故事是图像描述任务中的挑战,目标不仅是准确描述视觉内容,还需深入阐释作品意义。该任务对艺术图像尤其复杂,因不同艺术流派和风格存在多样化的解读方式与美学原则。为此,我们提出KALE(知识增强型视觉-语言模型用于艺术作品阐释),通过引入艺术元数据作为额外知识来增强现有视觉-语言模型。KALE以两种方式整合元数据:一是作为直接文本输入,二是通过多模态异构知识图谱。为优化图表示学习,我们设计了一种新的跨模态对齐损失,最大化图像与其对应元数据之间的相似性。实验结果表明,KALE在多个艺术图像数据集上表现优异,特别是在CIDEr指标上超越现有最先进方法。项目源码已公开于https://github.com/Yanbei-Jiang/Artwork-Interpretation。
原文摘要 · Abstract (English)
Exploring the narratives conveyed by fine-art paintings is a challenge in image captioning, where the goal is to generate descriptions that not only precisely represent the visual content but also offer a in-depth interpretation of the artwork's meaning. The task is particularly complex for artwork images due to their diverse interpretations and varied aesthetic principles across different artistic schools and styles. In response to this, we present KALE Knowledge-Augmented vision-Language model for artwork Elaborations), a novel approach that enhances existing vision-language models by integrating artwork metadata as additional knowledge. KALE incorporates the metadata in two ways: firstly as direct textual input, and secondly through a multimodal heterogeneous knowledge graph. To optimize the learning of graph representations, we introduce a new cross-modal alignment loss that maximizes the similarity between the image and its corresponding metadata. Experimental results demonstrate that KALE achieves strong performance (when evaluated with CIDEr, in particular) over existing state-of-the-art work across several artwork datasets. Source code of the project is available at https://github.com/Yanbei-Jiang/Artwork-Interpretation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。