AI模型a2z-1可精准识别腹部盆腔CT中的21种急症,性能稳定可靠。
a2z-1 for Multi-Disease Detection in Abdomen-Pelvis CT: External Validation and Performance Analysis Across 21 Conditions
- 基于深度学习构建多疾病检测模型,统一分析21类临床急症。
- 平均AUC达0.931,外部验证中仍保持0.923,关键病灶如肠梗阻达0.958。
- 跨性别、年龄、扫描参数均表现稳健,适合临床辅助诊断与质控。
我们对a2z-1人工智能模型在腹部盆腔CT中检测21种时间敏感且可干预的异常进行了全面评估。大规模回顾性分析显示,该模型在21种病症上的平均AUC为0.931。在两个不同医疗系统的外部验证中,模型表现稳定(AUC 0.923),在小肠梗阻(AUC 0.958)和急性胰腺炎(AUC 0.961)等危重病灶上尤为突出。亚组分析表明,模型在不同性别、年龄组及多种成像协议(包括不同层厚和对比剂使用方式)下均保持一致准确性。将高置信度输出与放射科报告对比,发现a2z-1能识别出被遗漏的异常,提示其在质量控制方面具有应用潜力。
原文摘要 · Abstract (English)
We present a comprehensive evaluation of a2z-1, an artificial intelligence (AI) model designed to analyze abdomen-pelvis CT scans for 21 time-sensitive and actionable findings. Our study focuses on rigorous assessment of the model's performance and generalizability. Large-scale retrospective analysis demonstrates an average AUC of 0.931 across 21 conditions. External validation across two distinct health systems confirms consistent performance (AUC 0.923), establishing generalizability to different evaluation scenarios, with notable performance in critical findings such as small bowel obstruction (AUC 0.958) and acute pancreatitis (AUC 0.961). Subgroup analysis shows consistent accuracy across patient sex, age groups, and varied imaging protocols, including different slice thicknesses and contrast administration types. Comparison of high-confidence model outputs to radiologist reports reveals instances where a2z-1 identified overlooked findings, suggesting potential for quality assurance applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。