arXiv:2605.17669cs.AI2026-05被引 1

用图文大模型扩展文化遗产知识图谱,提升信息整合效率与可靠性

Multimodal Cultural Heritage Knowledge Graph Extension with Language and Vision Models

论文配图:Multimodal Cultural Heritage Knowledge Graph Extension with Language and Vision Models
图 1 · 摘自论文原文
  • 结合图文大模型自动提取非结构化数据,实现多模态知识图谱扩展
  • 在法国文化遗产数据上构建了包含文本与图像的WJoconde图谱,支持知识补全任务
  • 开源完整数据集与工具链,适合文化遗产、AI多模态研究者使用

文化遗产的数字化保存与解读日益依赖数字技术,其中知识图谱(KG)因其结构化处理海量数据的能力而备受关注。然而,由于文化遗产信息的多样性与复杂性,知识图谱的构建与扩展面临挑战。本文提出一种面向法国文化遗产领域的新型知识图谱扩展方法,首先构建了多模态知识图谱WJoconde,集成实体的文本与图像信息,并推出三个变体以支持下游研究,如知识图谱补全(KGC)。同时,我们建立了针对该数据集的综合性KGC基准。其次,提出一种基于大语言模型(LLMs)和视觉-语言模型(VLMs)的多模态扩展框架,包含从非结构化资源中自动化提取数据,并通过专门的验证流程确保模型输出的准确性。实验表明,利用文化遗产中的丰富文本与图像信息,可高效且高可靠地扩展知识图谱。所有代码、基准数据集(含文本与图像)、原始数据及交互式访问接口均已开源。

原文摘要 · Abstract (English)

The preservation and interpretation of cultural heritage increasingly rely on digital technologies, among which Knowledge Graphs (KGs) stand out for their ability to structure vast amounts of data. However, the construction and expansion of these KGs often face challenges due to the diverse and complex nature of cultural heritage information. In this paper, we propose a novel approach for extending KG resources in the domain of cultural heritage, which we applied to French data. First, we introduce a new knowledge graph in the domain of French cultural heritage, WJoconde, which is distinguished by its multimodality as it integrates both textual and image information of the entities. We further introduce three variants of WJoconde to facilitate downstream research, such as Knowledge Graph Completion (KGC). We also built a comprehensive benchmark for KGC methods on our dataset. Second, we propose a new framework for extending cultural heritage KGs using multi-modal approaches leveraging Large Language Models (LLMs) and Vision-Language Models (VLMs), which includes automated data extraction from unstructured resources combined with a special validation pipeline for grounding the output of both models, to further extend WJoconde. Our results show that by integrating the rich text and image information in cultural heritage data, we can efficiently enhance KGs with high reliability. We open-source all code and benchmark datasets with text and images, as well as the original data with an interactive access point

知识图谱文化遗产多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。