对比五种NLP系统在儿科胸片报告标注中的表现,发现性能差异大,需谨慎评估。
Can Modern NLP Systems Reliably Annotate Chest Radiography Exams? A Pre-Purchase Evaluation and Comparative Study of Solutions from AWS, Google, Azure, John Snow Labs, and Open-Source Models on an Independent Pediatric Dataset
- 比较AWS、Google、Azure、SparkNLP及开源模型在儿科胸片报告中实体与诊断状态提取能力。
- 各系统提取实体数差异显著,准确率平均仅62%,最高76%(SparkNLP),最低50%(AWS)。
- 专用模型CheXpert/CheXbert准确率56%,提示通用工具不适用于特定临床任务。
通用临床自然语言处理(NLP)工具正被用于自动标注临床报告,但针对儿科胸片(CXR)报告标注等具体任务的独立评估仍有限。本研究对比了四种商用临床NLP系统——Amazon Comprehend Medical(AWS)、Google Healthcare NLP(GC)、Azure Clinical NLP(AZ)和SparkNLP(SP)——在儿科CXR报告中的实体抽取与断言检测能力。同时,使用CheXpert定义标签评估了专用模型CheXpert和CheXbert。分析来自大型学术儿童医院的95,008份儿科胸片报告,从发现与印象部分提取实体及其断言状态(阳性、阴性、不确定),并将印象部分实体映射至12类疾病及“无异常”类别。CheXpert与CheXbert均提取13类。以共识伪真实标签为基准,采用Fleiss Kappa与准确率进行对比。结果显示各系统提取实体数量差异显著:SP提取49,688个唯一实体,GC为16,477,AZ为31,543,AWS为27,216。断言准确率平均约62%,其中SP最高(76%),AWS最低(50%)。CheXpert与CheXbert准确率为56%。性能显著差异凸显了在部署前必须进行严格验证与审查。
原文摘要 · Abstract (English)
General-purpose clinical natural language processing (NLP) tools are increasingly used for the automatic labeling of clinical reports. However, independent evaluations for specific tasks, such as pediatric chest radiograph (CXR) report labeling, are limited. This study compares four commercial clinical NLP systems - Amazon Comprehend Medical (AWS), Google Healthcare NLP (GC), Azure Clinical NLP (AZ), and SparkNLP (SP) - for entity extraction and assertion detection in pediatric CXR reports. Additionally, CheXpert and CheXbert, two dedicated chest radiograph report labelers, were evaluated on the same task using CheXpert-defined labels. We analyzed 95,008 pediatric CXR reports from a large academic pediatric hospital. Entities and assertion statuses (positive, negative, uncertain) from the findings and impression sections were extracted by the NLP systems, with impression section entities mapped to 12 disease categories and a No Findings category. CheXpert and CheXbert extracted the same 13 categories. Outputs were compared using Fleiss Kappa and accuracy against a consensus pseudo-ground truth. Significant differences were found in the number of extracted entities and assertion distributions across NLP systems. SP extracted 49,688 unique entities, GC 16,477, AZ 31,543, and AWS 27,216. Assertion accuracy across models averaged around 62%, with SP highest (76%) and AWS lowest (50%). CheXpert and CheXbert achieved 56% accuracy. Considerable variability in performance highlights the need for careful validation and review before deploying NLP tools for clinical report labeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。