arXiv:2409.13744cs.CLcs.AI2024-09被引 10

用简单检索器提升LLM在疾病表型归一化中的准确率

A Simplified Retriever to Improve Accuracy of Phenotype Normalizations by Large Language Models

  • 基于BioBERT嵌入直接搜索HPO,无需显式术语定义
  • 在OMIM临床摘要上将归一化准确率从62.3%提升至90.3%
  • 方法轻量高效,适用于其他生物医学术语归一化任务

大型语言模型(LLMs)在表型术语归一化任务中,通过引入基于术语定义的候选归一化建议检索器,实现了准确率提升。本文提出一种简化检索器,利用BioBERT生成的上下文词嵌入,在人类表型本体(HPO)中搜索候选匹配,无需依赖显式术语定义。在来自在线孟德尔遗传病(OMIM)临床综述的术语上测试发现,该方法使当前最先进的LLM归一化准确率从无增强时的62.3%提升至90.3%。该方法具有可推广性,为更复杂的检索方式提供了一种高效替代方案。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown improved accuracy in phenotype term normalization tasks when augmented with retrievers that suggest candidate normalizations based on term definitions. In this work, we introduce a simplified retriever that enhances LLM accuracy by searching the Human Phenotype Ontology (HPO) for candidate matches using contextual word embeddings from BioBERT without the need for explicit term definitions. Testing this method on terms derived from the clinical synopses of Online Mendelian Inheritance in Man (OMIM), we demonstrate that the normalization accuracy of a state-of-the-art LLM increases from a baseline of 62.3% without augmentation to 90.3% with retriever augmentation. This approach is potentially generalizable to other biomedical term normalization tasks and offers an efficient alternative to more complex retrieval methods.

表型归一化LLM增强生物医学检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。