arXiv:2512.17279cs.CV2025-12被引 1

一个AI模型搞定多种超声诊断任务,但跨中心效果仍有差距。

Diagnostic Performance of Universal-Learning Ultrasound AI Across Multiple Organs and Tasks: the UUSIC25 Challenge

  • 用单一深度学习模型处理多器官分类与分割任务,提升临床实用性。
  • 顶级模型在5个分割任务上平均准确率达85.4%,乳腺癌分型在新中心下降至50.8%。
  • 适合追求通用性、关注模型泛化能力的医疗AI研究者参考。

现代超声设备可全身成像,但现有AI工具仍局限于单一任务,导致流程整合困难。本研究通过UUSIC25挑战赛,评估单一通用深度学习模型在多器官分类与分割中的诊断精度、泛化能力与效率。训练数据来自12个来源(9个公开、3个私有),共11,644张图像;测试集为独立多中心私有数据,含2,479张图像,其中一中心数据完全未参与训练以检验泛化能力。结果表明,在15个有效算法中,最优模型SMART在5个分割任务上实现宏平均Dice相似系数(DSC)0.854,二分类任务AUC达0.766。其在胎儿头颅等结构分割表现优异(DSC: 0.942),但在乳腺癌分子分型任务上,内部测试AUC为0.571,外部未见中心骤降至0.508,显示领域偏移问题严重。结论:通用模型可在多任务中达到高精度与高效,但跨中心泛化仍是临床部署的关键挑战。

原文摘要 · Abstract (English)

IMPORTANCE: Modern ultrasound systems are universal diagnostic tools capable of imaging the entire body. However, current AI solutions remain fragmented into single-task tools. This critical gap between hardware versatility and software specificity limits workflow integration and clinical utility. OBJECTIVE: To evaluate the diagnostic accuracy, versatility, and efficiency of single general-purpose deep learning models for multi-organ classification and segmentation. DESIGN: The Universal UltraSound Image Challenge 2025 (UUSIC25) involved developing algorithms on 11,644 images aggregated from 12 sources (9 public, 3 private). Evaluation used an independent, multi-center private test set of 2,479 images, including data from a center completely unseen during training to assess generalization. OUTCOMES: Diagnostic performance (Dice Similarity Coefficient [DSC]; Area Under the Receiver Operating Characteristic Curve [AUC]) and computational efficiency (inference time, GPU memory). RESULTS: Of 15 valid algorithms, the top model (SMART) achieved a macro-averaged DSC of 0.854 across 5 segmentation tasks and AUC of 0.766 for binary classification. Models demonstrated high capability in anatomical segmentation (e.g., fetal head DSC: 0.942) but variability in complex diagnostic tasks subject to domain shift. Specifically, in breast cancer molecular subtyping, the top model's performance dropped from an AUC of 0.571 (internal) to 0.508 (unseen external center), highlighting the challenge of generalization. CONCLUSIONS: General-purpose AI models can achieve high accuracy and efficiency across multiple tasks using a single architecture. However, significant performance degradation on unseen data suggests domain generalization is critical for future clinical deployment.

超声AI通用模型泛化能力医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。