用大模型生成生命科学领域多层次、结构复杂的本体,解决传统方法难以覆盖全类的问题。
LLMs4Life: Large Language Models for Ontology Learning in Life Sciences
- 通过提示工程与本体重用增强模型对生命科学的领域推理能力。
- 在AquaDiva本体上验证,生成结果具备逻辑一致性与可扩展性。
- 适合需要构建高复杂度本体的研究者,尤其关注生物医学知识图谱者。
生命科学等复杂领域中的本体学习对当前大型语言模型(LLMs)构成重大挑战。现有模型受限于生成文本长度及领域适配不足,难以生成具有多层结构、丰富关联和全面类别覆盖的本体。为此,我们改进了NeOn-GPT管道,引入先进提示工程与本体重用技术,提升生成本体的领域推理能力与结构深度。以合作研究中心AquaDiva开发并使用的AquaDiva本体为案例,评估生成本体的逻辑一致性、完整性与可扩展性。实验表明,该方法在高度专业化的生命科学领域中具备可行性,有效解决了模型性能与可扩展性的长期瓶颈。
原文摘要 · Abstract (English)
Ontology learning in complex domains, such as life sciences, poses significant challenges for current Large Language Models (LLMs). Existing LLMs struggle to generate ontologies with multiple hierarchical levels, rich interconnections, and comprehensive class coverage due to constraints on the number of tokens they can generate and inadequate domain adaptation. To address these issues, we extend the NeOn-GPT pipeline for ontology learning using LLMs with advanced prompt engineering techniques and ontology reuse to enhance the generated ontologies' domain-specific reasoning and structural depth. Our work evaluates the capabilities of LLMs in ontology learning in the context of highly specialized and complex domains such as life science domains. To assess the logical consistency, completeness, and scalability of the generated ontologies, we use the AquaDiva ontology developed and used in the collaborative research center AquaDiva as a case study. Our evaluation shows the viability of LLMs for ontology learning in specialized domains, providing solutions to longstanding limitations in model performance and scalability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。