arXiv:2412.08900cs.CLcs.AI2024-12中稿 · AMIA Annual Sympos…被引 2

用AI从医学文献中自动提取癌症治疗关键信息,辅助精准诊疗决策。

AI-assisted Knowledge Discovery in Biomedical Literature to Support Decision-making in Precision Oncology

  • 采用BERT和大模型等NLP技术,自动识别文献中的基因、药物等实体
  • BioBERT在关系抽取上表现最佳,F1值达0.79,几乎捕捉所有关键关联
  • 适合临床科研人员快速获取最新靶向治疗知识,提升决策效率

为向癌症患者提供合适的靶向治疗,需综合分析肿瘤分子特征与患者临床信息,并结合现有知识及最新生物医学文献。本文评估了特定自然语言处理方案在支持从生物医学文献中发现知识方面的潜力。测试了两个来自BERT家族的模型、两个大语言模型以及PubTator 3.0在命名实体识别(NER)和关系抽取(RE)任务中的表现。PubTator 3.0与BioBERT在NER任务中表现最佳,最佳F1分数分别为0.93和0.89;而BioBERT在RE任务中优于所有其他方案,最佳F1分数达0.79,并在特定应用场景中成功识别出几乎全部实体提及及多数关系。

原文摘要 · Abstract (English)

The delivery of appropriate targeted therapies to cancer patients requires the complete analysis of the molecular profiling of tumors and the patient's clinical characteristics in the context of existing knowledge and recent findings described in biomedical literature and several other sources. We evaluated the potential contributions of specific natural language processing solutions to support knowledge discovery from biomedical literature. Two models from the Bidirectional Encoder Representations from Transformers (BERT) family, two Large Language Models, and PubTator 3.0 were tested for their ability to support the named entity recognition (NER) and the relation extraction (RE) tasks. PubTator 3.0 and the BioBERT model performed best in the NER task (best F1-score equal to 0.93 and 0.89, respectively), while BioBERT outperformed all other solutions in the RE task (best F1-score 0.79) and a specific use case it was applied to by recognizing nearly all entity mentions and most of the relations.

精准医疗NLP应用文献挖掘癌症研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。