arXiv:2603.25862cs.CLcs.AI2026-03

用AI从海量文本自动构建可解释的知识图谱,覆盖新闻、建筑、医疗三领域。

Methods for Knowledge Graph Construction from Text Collections: Development and Applications

  • 融合NLP与语义网技术,实现跨文本类型的知识图谱自动生成。
  • 在三大应用场景中构建了可分析的图谱数据,支持趋势与因果关系挖掘。
  • 成果包括定制算法、基准测试和可复用的领域知识图谱资源。

社会各领域正面临非结构化文本数据的爆炸式增长,涵盖新闻、社交媒体、开放获取学术文献、数字健康记录及患者药物评价等。这些数据的规模与多样性为提取可操作知识提供了前所未有的机遇,也带来了严峻挑战。要从中提取丰富语义知识,需具备可扩展且灵活的自动化方法,能适应不同文本类型与模式规范。唯有将信息抽取与语义网技术结合,才能构建出语义透明、可解释、可互操作的完整知识图谱。本论文探索自然语言处理、机器学习与生成式AI方法,在语义网最佳实践指导下,从大规模文本语料中自动构建知识图谱,应用于三个场景:全球新闻与社交媒体中数字化转型话语的分析;建筑、工程、施工与运营领域大量出版物的映射与趋势分析;以及从电子健康记录与患者撰写的药物评价中生成生物医学实体间的因果关系图。论文贡献包括基准评估结果、定制算法设计,以及以知识图谱形式创建的数据资源,及其上层数据分析成果。

原文摘要 · Abstract (English)

Virtually every sector of society is experiencing a dramatic growth in the volume of unstructured textual data that is generated and published, from news and social media online interactions, through open access scholarly communications and observational data in the form of digital health records and online drug reviews. The volume and variety of data across all this range of domains has created both unprecedented opportunities and pressing challenges for extracting actionable knowledge for several application scenarios. However, the extraction of rich semantic knowledge demands the deployment of scalable and flexible automatic methods adaptable across text genres and schema specifications. Moreover, the full potential of these data can only be unlocked by coupling information extraction methods with Semantic Web techniques for the construction of full-fledged Knowledge Graphs, that are semantically transparent, explainable by design and interoperable. In this thesis, we experiment with the application of Natural Language Processing, Machine Learning and Generative AI methods, powered by Semantic Web best practices, to the automatic construction of Knowledge Graphs from large text corpora, in three use case applications: the analysis of the Digital Transformation discourse in the global news and social media platforms; the mapping and trend analysis of recent research in the Architecture, Engineering, Construction and Operations domain from a large corpus of publications; the generation of causal relation graphs of biomedical entities from electronic health records and patient-authored drug reviews. The contributions of this thesis to the research community are in terms of benchmark evaluation results, the design of customized algorithms and the creation of data resources in the form of Knowledge Graphs, together with data analysis results built on top of them.

知识图谱NLP生成式AI数据挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。