arXiv:2502.09659cs.CLcs.AI2025-02被引 3

用大模型自动识别癌症疫苗佐剂名称,提升科研效率。

Cancer Vaccine Adjuvant Name Recognition from Biomedical Literature using Large Language Models

  • 用GPT-4o和Llama3.2在零样本与少样本下识别佐剂名
  • 在VAC数据集上达77.32%的F1分数,优于Llama3.2-3B约2%
  • 适用于罕见命名形式,适合生物医学文献挖掘者

佐剂是增强疫苗效力、提升免疫反应的关键化学物质。从不断增长的生物医学文献中自动识别癌症疫苗佐剂名称,对推进研究和改进免疫疗法至关重要。本研究探索使用大语言模型(LLMs)——GPT-4o和Llama 3.2——实现自动化识别。采用两个数据集:97条来自AdjuvareDB的临床试验记录,以及290篇经疫苗佐剂词典(VAC)标注的摘要。在零样本与少样本学习范式下,最多每提示使用4个示例,通过明确指向佐剂名称的提示,测试上下文信息(如物质或干预措施)的影响。输出结果经自动化与人工验证。结果显示,GPT-4o在所有情境下均达到100%精确率,且在引入干预信息后召回率与F1分数显著提升。在VAC数据集上,其最大F1得分为77.32%,较Llama-3.2-3B高出约2%;在AdjuvareDB数据集上,三样本提示结合干预信息时,F1得分为81.67%,远超Llama-3.2-3B最高值65.62%。结论表明,大模型在识别佐剂名称方面表现优异,包括罕见命名形式。研究强调了其在加速癌症疫苗开发中的潜力,未来将拓展至更多生物医学文献并提升模型泛化能力。

原文摘要 · Abstract (English)

Motivation: An adjuvant is a chemical incorporated into vaccines that enhances their efficacy by improving the immune response. Identifying adjuvant names from cancer vaccine studies is essential for furthering research and enhancing immunotherapies. However, the manual curation from the constantly expanding biomedical literature poses significant challenges. This study explores the automated recognition of vaccine adjuvant names using Large Language Models (LLMs), specifically Generative Pretrained Transformers (GPT) and Large Language Model Meta AI (Llama). Methods: We utilized two datasets: 97 clinical trial records from AdjuvareDB and 290 abstracts annotated with the Vaccine Adjuvant Compendium (VAC). GPT-4o and Llama 3.2 were employed in zero-shot and few-shot learning paradigms with up to four examples per prompt. Prompts explicitly targeted adjuvant names, testing the impact of contextual information such as substances or interventions. Outputs underwent automated and manual validation for accuracy and consistency. Results: GPT-4o attained 100% Precision across all situations while exhibiting notable improve in Recall and F1-scores, particularly with incorporating interventions. On the VAC dataset, GPT-4o achieved a maximum F1-score of 77.32% with interventions, surpassing Llama-3.2-3B by approximately 2%. On the AdjuvareDB dataset, GPT-4o reached an F1-score of 81.67% for three-shot prompting with interventions, surpassing Llama-3.2-3 B's maximum F1-score of 65.62%. Conclusion: Our findings demonstrate that LLMs excel at identifying adjuvant names, including rare variations of naming representation. This study emphasizes the capability of LLMs to enhance cancer vaccine development by efficiently extracting insights. Future work aims to broaden the framework to encompass various biomedical literature and enhance model generalizability across various vaccines and adjuvants.

大模型命名实体识别癌症疫苗生物医学文本挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。