arXiv:2410.15186cs.CLcs.AI2024-10被引 7

用预训练模型自动编码兽医病历,提升数据可用性。

Fine-tuning foundational models to code diagnoses from veterinary health records

  • 用13个预训练语言模型微调兽医病历文本,覆盖7739个诊断代码
  • 基于24万份标注病历,大模型在有限资源下仍达良好准确率
  • 推动动物与人类健康数据融合,助力跨物种医疗研究

兽医医疗记录是兽医及人畜共患病临床研究的重要数据资源,但受限于格式不统一和数据孤岛问题。采用标准化医学术语进行临床编码可提升记录质量,并促进与人类健康数据的互操作性。先前研究如DeepTag和VetTag已尝试用NLP技术(如LSTM和Transformer)从自由文本中推断部分SNOMED-CT诊断代码。本研究进一步拓展,涵盖科罗拉多州立大学兽医教学医院(CSU VTH)认可的全部7,739个不同SNOMED-CT诊断代码,并利用日益丰富的预训练语言模型(LMs)。研究对13个免费可用的预训练LM在包含246,473例人工标注的兽医患者就诊记录的自由文本上进行微调,结果优于以往工作。使用大量标注数据微调较大规模临床LM时表现最佳,但也表明在资源有限情况下,使用非临床LM也能获得可比效果。该研究为自动化编码提供了可行路径,提升了兽医电子健康记录(EHR)质量,支持动物与人类健康研究的数据整合与共享。

原文摘要 · Abstract (English)

Veterinary medical records represent a large data resource for application to veterinary and One Health clinical research efforts. Use of the data is limited by interoperability challenges including inconsistent data formats and data siloing. Clinical coding using standardized medical terminologies enhances the quality of medical records and facilitates their interoperability with veterinary and human health records from other sites. Previous studies, such as DeepTag and VetTag, evaluated the application of Natural Language Processing (NLP) to automate veterinary diagnosis coding, employing long short-term memory (LSTM) and transformer models to infer a subset of Systemized Nomenclature of Medicine - Clinical Terms (SNOMED-CT) diagnosis codes from free-text clinical notes. This study expands on these efforts by incorporating all 7,739 distinct SNOMED-CT diagnosis codes recognized by the Colorado State University (CSU) Veterinary Teaching Hospital (VTH) and by leveraging the increasing availability of pre-trained language models (LMs). 13 freely-available pre-trained LMs were fine-tuned on the free-text notes from 246,473 manually-coded veterinary patient visits included in the CSU VTH's electronic health records (EHRs), which resulted in superior performance relative to previous efforts. The most accurate results were obtained when expansive labeled data were used to fine-tune relatively large clinical LMs, but the study also showed that comparable results can be obtained using more limited resources and non-clinical LMs. The results of this study contribute to the improvement of the quality of veterinary EHRs by investigating accessible methods for automated coding and support both animal and human health research by paving the way for more integrated and comprehensive health databases that span species and institutions.

兽医AI临床编码自然语言处理数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。