构建大规模医学影像分割基准,检验AI在真实场景下的泛化能力。
Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?
- 基于全球76家医院的5195例训练数据与11家医院的5903例测试数据
- 14个算法团队参与,第三方独立评估,覆盖多种分布外场景
- 评估开源框架如MONAI、nnU-Net,推动医疗AI持续创新
如何有效评估AI性能?看似简单实则复杂:现有基准常存在测试集分布单一、规模小、指标过于简化、比较不公及短期成果压力等问题,导致基准表现优异未必代表真实应用成功。为此,我们提出Touchstone——一个涵盖9类腹部器官的大规模协作分割基准。该基准包含来自全球76家医院的5,195例训练CT扫描和来自11家额外医院的5,903例测试CT扫描,多样化的测试集显著提升结果统计效力,并严格评估算法在多种分布外场景下的表现。我们邀请14位19个AI算法的发明者参与训练,而我们的团队作为第三方独立评估其在三个测试集上的表现。此外,还评估了若干预存的AI框架——这些框架更具灵活性,可支持多种算法,包括NVIDIA的MONAI、DKFZ的nnU-Net及其他开源框架。我们致力于持续扩展该基准,以促进医疗领域AI算法的持续创新。
原文摘要 · Abstract (English)
How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks--which, differing from algorithms, are more flexible and can support different algorithms--including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。