用大模型自动提取技术文档知识,生成可定制的知识图谱。
Leveraging LLM for Automated Ontology Extraction and Knowledge Graph Generation
- 通过自适应迭代思维链算法,让大模型交互式生成符合用户需求的本体。
- 支持无缝接入Neo4j等数据库,实现对非结构化数据的灵活查询与分析。
- 适合需要构建领域知识图谱的研发团队或智能系统开发者。
从可靠性与可维护性(RAM)领域的大型复杂技术文档中提取相关且结构化的知识,传统方式耗时且易出错。本文提出OntoKGen,一个真实的本体抽取与知识图谱(KG)生成流水线。该系统利用大语言模型(LLMs),通过由自适应迭代思维链(CoT)算法驱动的交互式界面,确保本体抽取及后续知识图谱生成过程符合用户特定需求。尽管知识图谱生成基于确认后的本体,但并无唯一正确本体,因其本质依赖于用户偏好。OntoKGen在遵循最佳实践的基础上推荐本体,降低用户负担并揭示潜在洞察,同时赋予用户对最终本体的完全控制权。基于确认本体生成知识图谱后,OntoKGen支持与无模式、非关系型数据库如Neo4j的无缝集成,实现对多样化非结构化来源知识的灵活存储与检索,促进高级查询、分析与决策。此外,生成的知识图谱可作为未来集成到检索增强生成(RAG)系统的基础,显著提升领域专用智能应用的开发能力。
原文摘要 · Abstract (English)
Extracting relevant and structured knowledge from large, complex technical documents within the Reliability and Maintainability (RAM) domain is labor-intensive and prone to errors. Our work addresses this challenge by presenting OntoKGen, a genuine pipeline for ontology extraction and Knowledge Graph (KG) generation. OntoKGen leverages Large Language Models (LLMs) through an interactive user interface guided by our adaptive iterative Chain of Thought (CoT) algorithm to ensure that the ontology extraction process and, thus, KG generation align with user-specific requirements. Although KG generation follows a clear, structured path based on the confirmed ontology, there is no universally correct ontology as it is inherently based on the user's preferences. OntoKGen recommends an ontology grounded in best practices, minimizing user effort and providing valuable insights that may have been overlooked, all while giving the user complete control over the final ontology. Having generated the KG based on the confirmed ontology, OntoKGen enables seamless integration into schemeless, non-relational databases like Neo4j. This integration allows for flexible storage and retrieval of knowledge from diverse, unstructured sources, facilitating advanced querying, analysis, and decision-making. Moreover, the generated KG serves as a robust foundation for future integration into Retrieval Augmented Generation (RAG) systems, offering enhanced capabilities for developing domain-specific intelligent applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。