用NLP技术让古印度医书知识可查可懂
NLP-Driven Knowledge Extraction and Thematic Classification of Translated Ancient Indian Medical Texts
- 结合NER与BERTopic提取疾病、药材等关键实体和主题
- 构建Neo4j知识图谱,可视化医书中的概念关系
- 为中医研究者、数字人文学者提供古代医学新工具
古印度医书如《苏鲁塔本集》记载了大量关于疾病、治疗和外科技术的信息。然而其古老格式与复杂词汇阻碍了信息获取与系统整理。本研究利用自然语言处理技术,包括命名实体识别(NER)、BERTopic主题建模及Neo4j知识图谱构建,对译本进行概念提取、分类与可视化。通过主题建模识别医学核心议题,利用NER结构化识别疾病、疗法、研究人员与药用植物等实体。基于图数据库的网络分析实现概念间语义关系的表达,支持知识检索与数字保存。结果表明,图数据库、主题建模与实体识别共同推动阿育吠陀历史医学智慧的计算组织,弥合传统文本与现代数据驱动研究之间的鸿沟。该方法促进历史文本分析、医学信息学与数字人文的发展,使古印度医学智慧更易获取与理解。
原文摘要 · Abstract (English)
Ancient Indian medical texts like Sushruta Samhita have extensive information on diseases, treatments, and surgical techniques. Yet, their ancient format and use of intricate vocabulary pose difficulties in accessibility and systematic ordering. The research here utilizes Natural Language Processing (NLP) methods like Named Entity Recognition (NER), BERTopic modeling, and Knowledge Graph development in Neo4j to extract, categorize, and visualize important concepts based on translated versions. Thematic classification with BERTopic allows for the identification of the underlying medical topics, whereas NER supports the structured entity recognition of diseases, treatments, researchers, and medicinal plants. Graphbased network analysis with Neo4j also allows for the semantic representation of relationship among extracted entities, supporting knowledge retrieval and digital preservation. The findings illustrate how graph databases, topic modeling, and entity recognition facilitate the computational organization of Ayurveda's historical medical wisdom, closing the gap between the conventional texts and contemporary data-driven inquiry. The suggested method promotes historical text analysis, medical informatics, and digital humanities to make ancient Indian medical wisdom more accessible and understandable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。