构建首个苏格兰多中心脑部影像报告数据集,助力临床NLP模型泛化能力研究。
GS-BrainText: A Multi-Site Brain Imaging Report Dataset from Generation Scotland for Clinical Natural Language Processing Development and Validation
- 整合5个苏格兰医疗区8511份脑部影像报告,24种疾病标注共2431份。
- 跨机构性能差异显著:F1值在22.22至100之间,年龄组间也存在波动。
- 适合研究语言变异、诊断不确定性表达及数据特性对NLP影响的团队使用。
我们提出GS-BrainText,一个来自苏格兰世代队列的8,511份脑部放射科报告的精选数据集,其中2,431份标注了24种脑部疾病表型。该多中心数据集涵盖五个苏格兰NHS卫生局,平均年龄58岁,中位年龄53岁,具有广泛的年龄分布,为开发和评估可泛化的临床自然语言处理(NLP)算法提供了独特资源。专家团队采用注释方案进行标注,各卫生局双人标注比例为10%-100%,并实施严格质量控制。基于与注释方案同步开发的规则型NLP系统EdIE-R进行基准测试,发现不同卫生局(F1: 86.13–98.13)、疾病表型(F1: 22.22–100)和年龄组(F1: 87.01–98.13)间存在性能差异,凸显了现有NLP工具泛化能力的关键挑战。该数据集填补了英国临床文本资源的重要空白,为研究语言变体、诊断不确定性表达及数据特征对NLP性能的影响提供了宝贵资源。
原文摘要 · Abstract (English)
We present GS-BrainText, a curated dataset of 8,511 brain radiology reports from the Generation Scotland cohort, of which 2,431 are annotated for 24 brain disease phenotypes. This multi-site dataset spans five Scottish NHS health boards and includes broad age representation (mean age 58, median age 53), making it uniquely valuable for developing and evaluating generalisable clinical natural language processing (NLP) algorithms and tools. Expert annotations were performed by a multidisciplinary clinical team using an annotation schema, with 10-100% double annotation per NHS health board and rigorous quality assurance. Benchmark evaluation using EdIE-R, an existing rule-based NLP system developed in conjunction with the annotation schema, revealed some performance variation across health boards (F1: 86.13-98.13), phenotypes (F1: 22.22-100) and age groups (F1: 87.01-98.13), highlighting critical challenges in generalisation of NLP tools. The GS-BrainText dataset addresses a significant gap in available UK clinical text resources and provides a valuable resource for the study of linguistic variation, diagnostic uncertainty expression and the impact of data characteristics on NLP system performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。