无需重新标注,通过修改字典即可动态调整胸部X光报告标签体系。
Reconfigurable Radiology Labels Without Relabeling
- 将自由文本报告转为多标签矩阵,用字典编辑实现标签体系重构。
- MIMIC-CXR数据集重配置仅需196秒,成本比重新标注低99.7%。
- 可识别43%超出原标签体系的发现,适配临床与研究灵活需求。
公开的胸部X光(CXR)数据集通常采用固定的小规模标签体系(如CheXpert-14),但原始自由文本报告包含更多诊断信息——哪些信息重要取决于任务、机构和阅片人。我们提出一个管道:将自由文本报告转化为多标签矩阵,并通过修改字典而非重新推理来重构标签体系,实现无需重新标注。一次预处理后,对22.3万份报告的MIMIC-CXR数据集进行标签体系重构仅需196秒,无API成本;而使用Claude Opus 4.7完成等效重标注需花费6.6千美元。基于58标签分类体系,我们发现43%的胸部X光检查包含至少一项不在CheXpert-14中的发现。基于这些标签训练的图像探测器,在共享目标上达到与CheXpert-14相当的性能,同时在专家审核的长尾标签上取得0.78的AUROC。结果表明,放射学标注的工作单位应从重新标注数据集,转变为结构化后调整标签配置。
原文摘要 · Abstract (English)
Public chest-radiograph (CXR) datasets are typically released with small, fixed label schemas such as CheXpert-14. However, the underlying free-text reports describe far more findings -- and which findings matter depends on the task, site, and reader. We release a pipeline that converts free-text reports into multi-label matrices and then reconfigures the label schema through dictionary edits rather than new inference passes, i.e., without relabeling the corpus. After this one-time pass, reconfiguring MIMIC-CXR (223K reports) from cached annotations takes 196 seconds with no API cost, compared to \$6.6K for an equivalent relabeling pass with Claude Opus 4.7. Using a 58-label taxonomy, we show that 43\% of CXR studies contain at least one finding outside CheXpert-14. Image probes trained on these labels match CheXpert-14 probes on shared targets while also reaching 0.78 AUROC on expert-reviewed long-tail labels that CheXpert-14 cannot represent. These results suggest a different unit of work for radiology labeling: once reports are structured, the label schema becomes a configuration to edit, not a corpus to relabel.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。