用大模型检测胸部CT数据集中的报告与标签不一致,提升数据质量。
Large Language Model-Assisted Cleaning of Report-Derived Labels in a Large-Scale Chest CT Dataset
- 通过大模型分析放射科报告,自动提取并比对标签
- 整体标签一致率达96.4%,淋巴结肿大差异最明显
- 多模型投票结果最优,适合医学数据清洗研究者
目的:评估大语言模型(LLM)辅助标签清洗能否识别CT-RATE这一大规模公开胸部CT数据集中报告与标签的不一致。材料与方法:经报告去重后,共识别出24,446份独特放射科报告。因微软Azure AI Foundry内容安全过滤,12份报告被排除在主分析外,剩余24,434份报告涵盖439,812个标签实例,覆盖18种异常类别。使用结构化JSON输出,由GPT-5.4从报告文本生成二值标签,并与现有CT-RATE标签进行对比。不一致样本由放射科医生裁定。此外,随机抽取100份报告,人工标注作为参考标准,用于比较原始标签、单个大模型标签及多模型多数投票标签的表现。结果:GPT-5.4生成标签与CT-RATE标签总体一致性达96.4%,Cohen's kappa为0.884。淋巴结肿大类别一致性最低。在不一致案例审定中,放射科医生支持GPT-5.4标签的比例分别为74.2%(一般不一致)和91.9%(淋巴结肿大不一致)。与放射科医生标注参考相比,多模型多数投票标签在标签宏平均F1分数和Cohen's kappa上表现最佳。结论:大模型辅助标签清洗可识别出临床有意义的报告-标签不一致,有望推动公共影像数据集的规模化质量提升。清洗后的数据集将公开,以支持后续研究。
原文摘要 · Abstract (English)
Purpose: To evaluate whether large language model (LLM)-assisted label cleaning can identify label-report discordance in CT-RATE, a large-scale public chest CT dataset. Materials and Methods: After report-level deduplication, 24,446 unique radiology reports were identified. Twelve reports were excluded from the primary GPT-5.4 analysis because of Microsoft Azure AI Foundry content-safety filtering, leaving 24,434 reports and 439,812 label instances across 18 abnormality categories. GPT-5.4-derived binary labels were generated from report text using structured JSON output and compared with existing CT-RATE labels. Discordant instances were adjudicated by radiologists. In addition, 100 randomly sampled reports were manually annotated to compare CT-RATE labels, individual LLM-derived labels, and multi-LLM majority-vote labels against radiologist-annotated reference labels. Results: Overall agreement between GPT-5.4-derived and CT-RATE labels was 96.4%, with Cohen's kappa of 0.884. Lymphadenopathy showed the lowest agreement and kappa. In discordance review, radiologist adjudication supported GPT-5.4-derived labels in 72 of 97 (74.2%) general discordant instances and 91 of 99 (91.9%) targeted lymphadenopathy discordant instances. Against radiologist-annotated reference labels, multi-LLM majority-vote labels achieved the highest label-macro-averaged F1 score and Cohen's kappa. Conclusion: LLM-assisted label cleaning identified clinically meaningful label-report discordance in CT-RATE and may support scalable quality improvement of public imaging datasets. The cleaned dataset will be made publicly available to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。