arXiv:2508.10703cs.AI2025-08中稿 · publication in Wor…被引 5

用大模型生成定义提升医学本体匹配效果

GenOM: Ontology Matching with Description Generation and Large Language Model

  • 用大模型生成概念描述,增强语义表示
  • 在生物本体匹配任务中性能超越多数基线
  • 适合需要精准知识融合的生物医药研究

本体匹配(OM)在实现异构知识源间的语义互操作与集成方面起着关键作用,尤其在包含大量疾病与药物相关复杂概念的生物医学领域。本文提出GenOM,一种基于大语言模型(LLM)的本体对齐框架,通过生成文本定义来丰富本体概念的语义表征,利用嵌入模型检索对齐候选,并结合基于精确匹配的工具提升精度。在OAEI Bio-ML赛道上进行的大量实验表明,GenOM通常能取得具有竞争力的表现,超越多种传统本体匹配系统及近期基于LLM的方法。进一步的消融实验验证了语义增强与少样本提示的有效性,凸显该框架的鲁棒性与适应性。

原文摘要 · Abstract (English)

Ontology matching (OM) plays an essential role in enabling semantic interoperability and integration across heterogeneous knowledge sources, particularly in the biomedical domain which contains numerous complex concepts related to diseases and pharmaceuticals. This paper introduces GenOM, a large language model (LLM)-based ontology alignment framework, which enriches the semantic representations of ontology concepts via generating textual definitions, retrieves alignment candidates with an embedding model, and incorporates exact matching-based tools to improve precision. Extensive experiments conducted on the OAEI Bio-ML track demonstrate that GenOM can often achieve competitive performance, surpassing many baselines including traditional OM systems and recent LLM-based methods. Further ablation studies confirm the effectiveness of semantic enrichment and few-shot prompting, highlighting the framework's robustness and adaptability.

本体匹配大模型生物医学语义增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。