arXiv:2608.10084eess.IVcs.CV2026-08中稿 · ed

检查医学影像标签真实性,避免模型学错数据。

When Repository Labels Are Not Image-Level Truth: A Supervision Auditing Framework for Chest Radiograph AI

论文配图:When Repository Labels Are Not Image-Level Truth: A Supervision Auditing Framework for Chest Radiograph AI
图 1 · 摘自论文原文
  • 用专家标注对比数据集标签,找出不一致根源。
  • 仅1%的标签匹配专家判断,近半误判为无异常。
  • 适合想训练可信医疗AI的研究者使用。

公开的胸部X光数据集广泛用于训练医疗AI,但其标签通常从放射科报告中提取,未经图像层面验证。本文提出仓库监督审计(RSA)框架,在模型开发前将数据集标签与放射科医生评审的图像标注进行比对。以MIMIC-CXR中的心大为例,发现仓库标签与专家图像标注几乎无一致,仅识别出1%的专家确认病例。多数差异源于未提及而非明确否定,近半被标记为“无发现”的图像实际存在心大。基于专家修正后的数据集,DenseNet121模型测试ROC-AUC达0.853。结果表明,仓库标签可能无法真实反映图像内容,监督审计是构建可靠医疗影像AI的关键步骤。

原文摘要 · Abstract (English)

Public chest X-ray repositories are widely used to train medical AI systems, yet their labels are typically extracted from radiology reports rather than verified directly on images. As a result, repository labels are often treated as image-level ground truth without validating whether they reflect what is actually visible in the radiograph. We introduce Repository Supervision Auditing (RSA), a framework that evaluates repository-derived labels against expert image-level annotations before model development. Using cardiomegaly in MIMIC-CXR as a case study, RSA compares repository labels with radiologist-reviewed image annotations, characterizes disagreement sources, and builds a curated cohort for deployment-oriented evaluation. Repository-derived cardiomegaly labels showed near-zero agreement with expert image-level assessment, identifying only 1% of expert-confirmed cases. Most discrepancies resulted from non-mention rather than explicit report negation, with expert-confirmed cardiomegaly identified in nearly half of studies assigned a repository-derived No Finding label. Using the resulting expert-curated cohort, a DenseNet121 model achieved a test ROC-AUC of 0.853. These findings show that repository labels may not reliably represent image-level truth and highlight supervision auditing as a critical step for developing trustworthy medical imaging AI.

医疗AI标签审计胸部X光模型可信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。