arXiv:2507.19396cs.CL2025-07被引 1

用Transformer模型在荷兰临床文本中检测药物不良反应,效果优于传统方法。

Detection of Adverse Drug Events in Dutch clinical free text documents using Transformer Models: benchmark study

  • 采用多种Transformer模型处理荷兰重症患者病程记录和出院小结的实体与关系识别。
  • MedRoBERTa.nl模型在金标准和端到端任务中均表现最佳,宏平均F1达0.63。
  • 外部验证显示该模型可检出67%至74%含不良反应的出院记录,适合临床部署。

本研究基于102份标注丰富的荷兰重症监护室临床病程记录,对比了Bi-LSTM与四种荷兰/多语言Transformer编码器(BERTje、RobBERT、MedRoBERTa.nl、NuNER)在命名实体识别(NER)与关系分类(RC)任务中的表现。数据来自一所教学医院的重症监护室病程记录及两所非教学医院内科病房的出院小结。通过金标准(两步法)与预测实体(端到端)两种方式评估模型,并进行外部文档级不良药物事件(ADE)检测验证。报告微平均与宏平均F1分数以应对标签不平衡问题。尽管各模型差异较小,但MedRoBERTa.nl在金标准下取得0.63的宏平均F1,在预测实体下为0.62。外部验证中,该模型召回率达0.67至0.74,即成功识别67%至74%的含不良反应的出院记录。研究建立了针对临床自由文本中药品不良反应检测的可靠基准,强调需采用适配任务的评估指标,为未来临床应用提供支持。

原文摘要 · Abstract (English)

In this study, we establish a benchmark for adverse drug event (ADE) detection in Dutch clinical free-text documents using several transformer models, clinical scenarios, and fit-for-purpose performance measures. We trained a Bidirectional Long Short-Term Memory (Bi-LSTM) model and four transformer-based Dutch and/or multilingual encoder models (BERTje, RobBERT, MedRoBERTa(.)nl, and NuNER) for the tasks of named entity recognition (NER) and relation classification (RC) using 102 richly annotated Dutch ICU clinical progress notes. Anonymized free-text clinical progress notes of patients admitted to the intensive care unit (ICU) of one academic hospital and discharge letters of patients admitted to Internal Medicine wards of two non-academic hospitals were reused. We evaluated our ADE RC models internally using the gold standard (two-step task) and predicted entities (end-to-end task). In addition, all models were externally validated for detecting ADEs at the document level. We report both micro- and macro-averaged F1 scores, given the dataset imbalance in ADEs. Although differences for the ADE RC task between the models were small, MedRoBERTa(.)nl was the best performing model with a macro-averaged F1 score of 0.63 using the gold standard and 0.62 using predicted entities. The MedRoBERTa(.)nl models also performed the best in our external validation and achieved a recall of between 0.67 to 0.74 using predicted entities, meaning between 67 to 74% of discharge letters with ADEs were detected. Our benchmark study presents a robust and clinically meaningful approach for evaluating language models for ADE detection in clinical free-text documents. Our study highlights the need to use appropriate performance measures fit for the task of ADE detection in clinical free-text documents and envisioned future clinical use.

药物安全自然语言处理临床文本挖掘Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。