用大模型把文化遗产文本转成可查询的知识图谱,让争议性内容也能结构化。
Knowledge Graphs Generation from Cultural Heritage Texts: Combining LLMs and Ontological Engineering for Scholarly Debates
- 五步法融合大模型与本体工程,逐层构建文化遗产知识图谱
- 提取准确率达0.99(元数据),证据抽取达0.97,小模型表现不输大模型
- 适合文化机构做知识系统化,尤其适合有争议文献的自动化处理
文化遗产文本蕴含丰富知识,但因难以将非结构化论述转化为结构化知识图谱(KG),查询困难。本文提出ATR4CH(自适应文本到RDF的文化遗产框架),一种基于大模型的知识提取系统性五步方法。通过真伪争议议题案例验证:包含基础分析、标注模式设计、流水线架构、集成优化与综合评估。使用三款大模型(Claude Sonnet 3.7、Llama 3.3 70B、GPT-4o-mini)处理维基百科中关于争议性文物/文件的文章。结果显示:元数据提取F1为0.96–0.99,实体识别0.7–0.8,假设抽取0.65–0.75,证据抽取0.95–0.97,话语表征得分0.62 G-EVAL。小模型表现良好,具备成本效益。这是首个将大模型提取与文化遗产本体协同的系统性方法,可跨领域复用。局限在于仅基于维基百科,需人工后处理。实际应用上,该方法助力文博机构将文本知识转化为可检索的知识图谱,支持元数据自动增强与知识发现。
原文摘要 · Abstract (English)
Cultural Heritage texts contain rich knowledge that is difficult to query systematically due to the challenges of converting unstructured discourse into structured Knowledge Graphs (KGs). This paper introduces ATR4CH (Adaptive Text-to-RDF for Cultural Heritage), a systematic five-step methodology for Large Language Model-based Knowledge Extraction from Cultural Heritage documents. We validate the methodology through a case study on authenticity assessment debates. Methodology - ATR4CH combines annotation models, ontological frameworks, and LLM-based extraction through iterative development: foundational analysis, annotation schema development, pipeline architecture, integration refinement, and comprehensive evaluation. We demonstrate the approach using Wikipedia articles about disputed items (documents, artifacts...), implementing a sequential pipeline with three LLMs (Claude Sonnet 3.7, Llama 3.3 70B, GPT-4o-mini). Findings - The methodology successfully extracts complex Cultural Heritage knowledge: 0.96-0.99 F1 for metadata extraction, 0.7-0.8 F1 for entity recognition, 0.65-0.75 F1 for hypothesis extraction, 0.95-0.97 for evidence extraction, and 0.62 G-EVAL for discourse representation. Smaller models performed competitively, enabling cost-effective deployment. Originality - This is the first systematic methodology for coordinating LLM-based extraction with Cultural Heritage ontologies. ATR4CH provides a replicable framework adaptable across CH domains and institutional resources. Research Limitations - The produced KG is limited to Wikipedia articles. While the results are encouraging, human oversight is necessary during post-processing. Practical Implications - ATR4CH enables Cultural Heritage institutions to systematically convert textual knowledge into queryable KGs, supporting automated metadata enrichment and knowledge discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。