面向罕见病与新病种的多中心胸部X光分类挑战,推动医学影像泛化能力研究。
Overview of the CXR-LT 2026 Challenge: Multi-Center Long-Tailed and Zero Shot Chest X-ray Classification
- 构建包含14.5万张图像的多中心数据集,模拟真实临床长尾分布。
- 任务一mAP达0.5854,任务二零样本识别mAP达0.4315,验证视觉语言模型优势。
- 适合关注医学AI泛化性、罕见病检测的研究者与临床工程师。
胸部X光(CXR)解读受限于病种分布的长尾特性及临床环境的开放世界本质。现有基准多依赖单一机构的封闭类别,难以反映罕见疾病的真实分布或新出现的病变。为此,我们推出CXR-LT 2026挑战赛。作为该基准的第三届,它引入了来自PadChest与NIH胸部X光数据集的多中心数据,涵盖超过145,000张图像。挑战设置两个核心任务:(1) 在30个已知类别上进行鲁棒的多标签分类;(2) 面向6个未见(分布外)的罕见病类别实现开放世界泛化。我们报告了顶尖团队的表现,评估指标包括平均精度(mAP)、AUROC和F1分数。优胜方案在任务一中取得mAP 0.5854,任务二中达到mAP 0.4315,表明大规模视觉-语言预训练能显著缓解零样本诊断中的性能下降问题。
原文摘要 · Abstract (English)
Chest X-ray (CXR) interpretation is hindered by the long-tailed distribution of pathologies and the open-world nature of clinical environments. Existing benchmarks often rely on closed-set classes from single institutions, failing to capture the prevalence of rare diseases or the appearance of novel findings. To address this, we present the CXR-LT 2026 challenge. This third iteration of the benchmark introduces a multi-center dataset comprising over 145,000 images from PadChest and NIH Chest X-ray datasets. The challenge defines two core tasks: (1) Robust Multi-Label Classification on 30 known classes and (2) Open-World Generalization to 6 unseen (out-of-distribution) rare disease classes. We report the results of the top-performing teams, evaluating them via mean Average Precision (mAP), AUROC, and F1-score. The winning solutions achieved an mAP of 0.5854 on Task 1 and 0.4315 on Task 2, demonstrating that large-scale vision-language pre-training significantly mitigates the performance drop typically associated with zero-shot diagnosis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。