用新数据集提升俄语病历编码自动化,效果优于人工标注。
RuCCoD: Towards Automated ICD Coding in Russian
- 构建俄语临床编码数据集,含超万条实体与1500+唯一ICD码。
- 自动化预测代码训练后,准确率显著高于医生人工标注。
- 适合关注低资源语言医疗自动化、临床数据标准化的研究者。
本研究探索在俄语这一生物医学资源有限的语言中实现临床编码自动化的可行性。我们构建了一个新的用于ICD编码的数据集,包含来自电子健康记录(EHR)的诊断字段,共标注超过10,000个实体和1,500多个唯一ICD代码。该数据集作为多个前沿模型(包括BERT、LLaMA+LoRA、RAG)的基准测试平台,并进一步考察了跨领域(从PubMed摘要到医学诊断)和跨术语体系(从UMLS概念到ICD代码)的迁移学习效果。随后,我们将表现最佳的模型应用于自建的内部EHR数据集(2017–2021年患者病史),实验基于精心筛选的测试集进行,结果表明:使用自动预测的代码进行训练,可显著提升模型准确率,优于基于医生人工标注的数据。研究结果为资源匮乏语言如俄语中的临床编码自动化提供了重要启示,有助于提升临床效率与数据准确性。代码与数据集已公开于 https://github.com/auto-icd-coding/ruccod。
原文摘要 · Abstract (English)
This study investigates the feasibility of automating clinical coding in Russian, a language with limited biomedical resources. We present a new dataset for ICD coding, which includes diagnosis fields from electronic health records (EHRs) annotated with over 10,000 entities and more than 1,500 unique ICD codes. This dataset serves as a benchmark for several state-of-the-art models, including BERT, LLaMA with LoRA, and RAG, with additional experiments examining transfer learning across domains (from PubMed abstracts to medical diagnosis) and terminologies (from UMLS concepts to ICD codes). We then apply the best-performing model to label an in-house EHR dataset containing patient histories from 2017 to 2021. Our experiments, conducted on a carefully curated test set, demonstrate that training with the automated predicted codes leads to a significant improvement in accuracy compared to manually annotated data from physicians. We believe our findings offer valuable insights into the potential for automating clinical coding in resource-limited languages like Russian, which could enhance clinical efficiency and data accuracy in these contexts. Our code and dataset are available at https://github.com/auto-icd-coding/ruccod.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。