arXiv:2412.03176cs.CL2024-12

用大模型与医学本体结合,自动识别西班牙语病历中的皮肤病

Automatic detection of diseases in Spanish clinical notes combining medical language models and ontologies

  • 结合语言模型与医学本体,按部位、严重程度、顺序学习病理特征
  • 在西班牙语病历上达到0.84精确率,0.82微F1,0.75宏F1
  • 方法与数据集开源,适合医疗文本分析与多语言NLP研究者

本文提出一种混合方法,用于自动检测医学报告中的皮肤疾病。通过将大语言模型与医学本体结合,根据首次就诊或复诊报告预测患者可能患有的病理类型。实验表明,引导模型按部位、严重程度和学习顺序来理解病理特征,可显著提升准确率。该方法在医疗文本分类任务中达到0.84的精确率、0.82的微F1得分和0.75的宏F1得分,性能达当前最优水平。论文同时公开了所用方法与数据集,供社区使用。

原文摘要 · Abstract (English)

In this paper we present a hybrid method for the automatic detection of dermatological pathologies in medical reports. We use a large language model combined with medical ontologies to predict, given a first appointment or follow-up medical report, the pathology a person may suffer from. The results show that teaching the model to learn the type, severity and location on the body of a dermatological pathology, as well as in which order it has to learn these three features, significantly increases its accuracy. The article presents the demonstration of state-of-the-art results for classification of medical texts with a precision of 0.84, micro and macro F1-score of 0.82 and 0.75, and makes both the method and the data set used available to the community.

皮肤病检测医疗NLP大模型应用多语言文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。