arXiv:2510.20727cs.CLcs.AI2025-10被引 2

用大模型从病历中自动提取氟尿嘧啶治疗与副作用信息

Automated Extraction of Fluoropyrimidine Treatment and Treatment-Related Toxicities from Clinical Notes Using Natural Language Processing

  • 采用大模型零样本和错误分析提示,提升医学文本信息抽取精度
  • 错误分析提示法对治疗和毒性提取的F1达1.000,显著优于其他方法
  • 适合临床研究、药物安全监测等需要高效挖掘病历数据的场景

氟尿嘧啶类药物广泛用于结直肠癌和乳腺癌治疗,但常伴随手足综合征、心毒性等副作用。由于这些信息多嵌入于临床笔记中,我们开发并评估了自然语言处理(NLP)方法以提取治疗方案与毒性信息。基于204,165名成年肿瘤患者共236份临床笔记构建了金标准数据集,由领域专家标注治疗方案与毒性类别。对比规则、机器学习(随机森林、支持向量机、逻辑回归)、深度学习(BERT、ClinicalBERT)及大语言模型(LLM)方法(零样本与错误分析提示)。使用80:20划分训练测试集。结果显示,足够数据支持5个类别的训练与评估。错误分析提示法在治疗与毒性提取上均取得最优结果(F1=1.000),零样本提示法治疗提取F1=1.000,毒性提取F1=0.876;逻辑回归与支持向量机在毒性提取中表现次优(F1=0.937)。深度学习方法表现较差,BERT(治疗F1=0.873,毒性F1=0.839)、ClinicalBERT(治疗F1=0.873,毒性F1=0.886)均低于其他方法。规则方法为基线,治疗与毒性提取的F1分别为0.857和0.858。讨论表明,基于大语言模型的方法优于所有其他方法,其次为机器学习方法。机器与深度学习受限于小样本数据,泛化能力差,尤其对罕见类别。结论:大语言模型最有效地从临床笔记中提取氟尿嘧啶治疗与毒性信息,具有支持肿瘤研究与药物警戒的强大潜力。

原文摘要 · Abstract (English)

Objective: Fluoropyrimidines are widely prescribed for colorectal and breast cancers, but are associated with toxicities such as hand-foot syndrome and cardiotoxicity. Since toxicity documentation is often embedded in clinical notes, we aimed to develop and evaluate natural language processing (NLP) methods to extract treatment and toxicity information. Materials and Methods: We constructed a gold-standard dataset of 236 clinical notes from 204,165 adult oncology patients. Domain experts annotated categories related to treatment regimens and toxicities. We developed rule-based, machine learning-based (Random Forest, Support Vector Machine [SVM], Logistic Regression [LR]), deep learning-based (BERT, ClinicalBERT), and large language models (LLM)-based NLP approaches (zero-shot and error-analysis prompting). Models used an 80:20 train-test split. Results: Sufficient data existed to train and evaluate 5 annotated categories. Error-analysis prompting achieved optimal precision, recall, and F1 scores (F1=1.000) for treatment and toxicities extraction, whereas zero-shot prompting reached F1=1.000 for treatment and F1=0.876 for toxicities extraction.LR and SVM ranked second for toxicities (F1=0.937). Deep learning underperformed, with BERT (F1=0.873 treatment; F1= 0.839 toxicities) and ClinicalBERT (F1=0.873 treatment; F1 = 0.886 toxicities). Rule-based methods served as our baseline with F1 scores of 0.857 in treatment and 0.858 in toxicities. Discussion: LMM-based approaches outperformed all others, followed by machine learning methods. Machine and deep learning approaches were limited by small training data and showed limited generalizability, particularly for rare categories. Conclusion: LLM-based NLP most effectively extracted fluoropyrimidine treatment and toxicity information from clinical notes, and has strong potential to support oncology research and pharmacovigilance.

医疗NLP药物安全大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。