面向罕见病与未知病灶的多中心胸部X光分类挑战,推动AI在真实临床环境中的应用。
CXR-LT 2026 Challenge: Multi-Center Long-Tailed and Zero Shot Chest X-ray Classification

- 构建超过14.5万张多中心胸部X光数据集,涵盖30种已知病种和6种未见病种。
- 首次采用放射科医生标注,提升标签可靠性,验证视觉-语言模型在罕见病识别中的优势。
- 聚焦真实场景下的长尾分布与开放世界泛化,适合医疗AI研发与评估人员参考。
胸部X光(CXR)诊断受病理长尾分布和临床环境开放性影响。现有基准多基于单一机构的封闭类别,难以反映罕见疾病普遍性或新出现病灶。为此,我们推出CXR-LT 2026挑战赛,整合来自PadChest和NIH Chest X-ray数据集的超14.5万张图像,所有开发与测试集均由放射科医生标注,确保临床可靠性。今年设定两大任务:(1)在30个已知类别上进行鲁棒多标签分类;(2)对6个未见(分布外)罕见病种实现开放世界泛化。本文总结挑战赛整体设计,分析参赛团队策略,评估头尾性能、校准性及跨中心泛化差距。结果显示,视觉-语言基础模型显著提升分布内与零样本表现,但面对多中心数据偏移时仍难有效检测罕见病灶。本研究为真实临床条件下AI系统的开发与评估提供坚实基础。
原文摘要 · Abstract (English)
Chest X-ray (CXR) interpretation is hindered by the long-tailed distribution of pathologies and the open-world nature of clinical environments. Existing benchmarks often rely on closed-set classes from a single institution, failing to capture the prevalence of rare diseases or the appearance of novel findings. To address this, we present the CXR-LT challenge. The first event, CXR-LT 2023, established a large-scale benchmark for long-tailed multi-label CXR classification and identified key challenges in rare disease recognition. CXR-LT 2024 further expanded the label space and introduced a zero-shot task to study generalization to unseen findings. Building on the success of CXR-LT 2023 and 2024, this third iteration of the benchmark introduces a multi-center dataset comprising over 145,000 images from PadChest and NIH Chest X-ray datasets. Additionally, all development and test sets in CXR-LT 2026 are annotated by radiologists, providing a more reliable and clinically grounded evaluation than report-derived labels. The challenge defines two core tasks this year: (1) Robust Multi-Label Classification on 30 known classes and (2) Open-World Generalization to 6 unseen (out-of-distribution) rare disease classes. This paper summarizes the overview of the CXR-LT 2026 challenge. We describe the data collection and annotation procedures, analyze solution strategies adopted by participating teams, and evaluate head-versus-tail performance, calibration, and cross-center generalization gaps. Our results show that vision-language foundation models improve both in-distribution and zero-shot performance, but detecting rare findings under multi-center shift remains challenging. Our study provides a foundation for developing and evaluating AI systems in realistic long-tailed and open-world clinical conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。