构建首个大规模意大利语急诊病历数据集,支持医疗大模型研究。
eCREAM-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

- 收集400万条匿名急诊病历,覆盖患者全流程诊疗。
- 6000份病历经专家标注132项指标,含数值、类别等多类型数据。
- 适合医疗AI研究者,尤其关注临床语言模型与结构化信息抽取。
本文介绍eCREAM-MedCorpus,一个来自意大利医院急诊科的全新大规模临床笔记数据集。当前版本包含约400万条完全匿名的临床笔记,涵盖患者在急诊科停留期间的多个阶段。此外,约6000条笔记由临床专家通过结构化病例报告表(CRF)手动标注,包含132个与急症场景(呼吸困难和意识丧失)相关的关键项目,数据类型包括数值型(如血氧饱和度)、类别型(如意识水平)、二值型(如是否存在创伤)及混合类型。标注过程由多位临床医生参与,经多次迭代修订以解决条目歧义,形成结构丰富但存在高不平衡性的资源。该数据集旨在填补意大利语临床数据的空白,支持大语言模型在真实医疗场景中的开发与应用。我们详细描述了数据采集流程、现场去匿名化管道、语料库统计信息及标注方案。最后,提出基于CRF填空的新型结构化信息抽取基准,并提供Gemma-27B与MedGemma-27B的零样本基线结果。据我们所知,eCREAM-MedCorpus是目前公开可用的最大意大利语临床笔记数据集。
原文摘要 · Abstract (English)
We present eCREAM-MedCorpus, a new and unique large-scale dataset of clinical notes produced in Emergency Departments of Italian hospitals. The corpus, in its current version, is composed of approximately 4 million clinical notes fully anonymized, covering diverse phases of patient care during the stay in the emergency department. In addition, a subset of about six thousand notes has been manually annotated by clinical experts through a structured Case Report Form (CRF) containing 132 items relevant for two patient situations in emergency departments, dyspnea and loss of consciousness. Items may assume numerical values (e.g., for blood saturation), categorical (e.g., for level of consciousness ), binary (e.g., for presence of traumas), and mixed value types. The annotation process involved multiple clinicians and underwent iterative revision to resolve ambiguities in item formulation, resulting in a richly structured (although high imbalanced) resource. The dataset aims to fill a relevant gap of data able to support both the development and the use of Large Language Models in concrete medical applications. We describe the data collection protocol, the on-site anonymisation pipeline, corpus statistics, and the annotation scheme. Finally, we propose CRF-filling as a novel structured information extraction benchmark, and provide zero-shot baseline resulting from Gemma-27B and MedGemma-27B. To the best of our knowledge, eCREAM-MedCorpus is the largest freely available dataset of clinical notes existing for the Italian language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。