构建首个覆盖8.5万例的癌症检测AI评估基准,揭示模型在少数群体中表现严重下降
BenchX: Benchmarking AI Models for Cancer Detection and Localization with Demographic and Protocol Biases

- 基于8.5万例CT数据,系统评估12种肿瘤检测模型在不同人群与扫描条件下的表现
- 顶尖模型在年轻女性非裔患者等少数群体中准确率显著下降,最低不足40%
- 利用大语言模型自动提取临床信息,实现可扩展、可复现的子群分析
人工智能在医学影像领域取得显著进展,但其在真实临床环境中表现不一,尤其当患者人口统计特征和成像协议存在差异时,如小肿瘤检测、不同增强相位图像分析或不同年龄/性别患者评估。为量化此类不一致性,我们构建了一个大规模开源基准,包含85,355例CT扫描,系统评估12种肿瘤检测AI模型在肿瘤大小、位置、患者子群及成像协议上的表现。我们利用大语言模型(LLMs)从临床数据中提取并组织子群信息,使分析具备可扩展性和可复现性。结果显示,当前最先进模型虽在平均准确率上表现优异,但在罕见或代表性不足的子群中表现不佳,例如年轻女性非裔群体,其准确率最低低于40%。然而,针对这些罕见病例收集足够标注数据往往不切实际。该基准为构建更可靠、鲁棒的肿瘤检测AI模型奠定基础,并强调在医学影像与计算机视觉中开展严格的子群层面评估的必要性。数据集与代码已公开。
原文摘要 · Abstract (English)
Artificial intelligence (AI) has achieved remarkable success in medical imaging, but it is widely recognized that these models often perform inconsistently across real-world clinical settings. Such inconsistencies occur when patient demographics and imaging protocols vary, for example, in detecting small tumors, analyzing scans from different contrast phases, or evaluating patients of different ages or sexes. To quantify these inconsistencies, we develop a large-scale, open benchmark of 85,355 CT scans that systematically evaluates 12 tumor-detection AI models across tumor size, location, patient subgroup, and imaging protocol. We leverage large language models (LLMs) to extract and organize subgroup information from clinical data, which makes the analysis both scalable and reproducible. Our benchmark reveals that current state-of-the-art AI models, optimized for average accuracy, perform poorly in rare or underrepresented subgroups, such as young, female African Americans. However, collecting sufficient annotated data for these rare cases is often impractical. The benchmark provides a foundation for building more reliable and robust AI models for tumor detection and highlighting the need for rigorous, subgroup-level evaluation in medical imaging and computer vision. Datasets, code
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。