LLM映射生物医学术语时,词频越高越准,低频词容易出错。
Mapping Biomedical Ontology Terms to IDs: Effect of Domain Prevalence on Prediction Accuracy
- 用文献中术语出现频次预测大模型映射准确率,发现频次越高越准。
- 高频术语映射准确率超95%,低频术语则明显下降。
- 提示训练数据需考虑术语流行度,尤其关注罕见词的识别。
本研究评估大型语言模型(LLMs)在人类表型本体(HPO)、基因本体(GO)和UniProtKB术语中,将生物医学术语映射到对应标识符(ID)的能力。以PubMed Central(PMC)数据集中术语ID的出现次数作为其在生物医学文献中流行度的代理指标,分析了术语流行度与映射准确率的关系。结果表明,术语流行度能有效预测HPO、GO术语及蛋白质名称对相应ID的映射准确性:文献中出现频率越高,映射准确率越高。基于受试者工作特征(ROC)曲线的预测模型验证了这一关系。然而,该模式不适用于蛋白质名称到人类基因组织(HUGO)基因符号的映射——GPT-4在此任务上已达到95%的高基线性能,且准确率不受流行度影响。研究认为,由于HUGO基因符号在文献中高度流行,已形成词汇化特征,使GPT-4可稳定映射。研究揭示了大模型在低流行度术语映射上的局限性,并强调在生物医学应用中应将术语流行度纳入训练与评估体系。
原文摘要 · Abstract (English)
This study evaluates the ability of large language models (LLMs) to map biomedical ontology terms to their corresponding ontology IDs across the Human Phenotype Ontology (HPO), Gene Ontology (GO), and UniProtKB terminologies. Using counts of ontology IDs in the PubMed Central (PMC) dataset as a surrogate for their prevalence in the biomedical literature, we examined the relationship between ontology ID prevalence and mapping accuracy. Results indicate that ontology ID prevalence strongly predicts accurate mapping of HPO terms to HPO IDs, GO terms to GO IDs, and protein names to UniProtKB accession numbers. Higher prevalence of ontology IDs in the biomedical literature correlated with higher mapping accuracy. Predictive models based on receiver operating characteristic (ROC) curves confirmed this relationship. In contrast, this pattern did not apply to mapping protein names to Human Genome Organisation's (HUGO) gene symbols. GPT-4 achieved a high baseline performance (95%) in mapping protein names to HUGO gene symbols, with mapping accuracy unaffected by prevalence. We propose that the high prevalence of HUGO gene symbols in the literature has caused these symbols to become lexicalized, enabling GPT-4 to map protein names to HUGO gene symbols with high accuracy. These findings highlight the limitations of LLMs in mapping ontology terms to low-prevalence ontology IDs and underscore the importance of incorporating ontology ID prevalence into the training and evaluation of LLMs for biomedical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。