arXiv:2603.26785eess.IVcs.CV2026-03被引 1

提出可复现的AI肺结节检测部署后评估框架,发现切片厚度比噪声更影响模型性能。

Beyond Benchmarks: A Framework for Post Deployment Validation of CT Lung Nodule Detection AI

  • 基于物理模拟生成五种扫描参数变化,评估模型敏感性
  • 5mm切片厚度使检出率下降至26.2%,比基线降低42%
  • 无需专有设备数据,适合资源受限环境持续质量监控

背景:人工智能辅助肺结节检测系统在临床部署时缺乏站点特异性验证。基准条件下的表现可能无法反映实际应用中因扫描参数差异导致的真实行为。目的:提出并验证一种基于物理引导的框架,用于评估已部署肺结节检测模型对CT扫描参数系统性变化的敏感性。方法:使用公开的LIDC-IDRI数据集中的21例,采用在LUNA16上预训练(折0,未微调)的MONAI RetinaNet模型,在五种成像条件下测试:基线、25%剂量降低、50%剂量降低、3mm层厚、5mm层厚。剂量降低通过图像域高斯噪声模拟;层厚变化通过z轴方向移动平均实现。检测敏感性在置信度阈值0.5下计算,匹配标准为15mm。结果:基线敏感性为45.2%(57/126个共识结节)。剂量降低引起轻微退化:25%剂量下为41.3%,50%剂量下为42.1%。5mm层厚条件下敏感性显著下降至26.2%,相比基线下降19个百分点,相对减少42%。该现象在0.1至0.9的置信度阈值范围内均一致。病例级分析显示性能异质性明显,其中两例在基线即出现完全漏检。结论:在所测试条件下,层厚是比图像噪声更根本的制约因素。所提框架具有可复现性,无需专有扫描仪数据,适用于资源受限环境下持续的部署后质量保证。

原文摘要 · Abstract (English)

Background: Artificial intelligence (AI) assisted lung nodule detection systems are increasingly deployed in clinical settings without site-specific validation. Performance reported under benchmark conditions may not reflect real-world behavior when acquisition parameters differ from training data. Purpose: To propose and demonstrate a physics-guided framework for evaluating the sensitivity of a deployed lung nodule detection model to systematic variation in CT acquisition parameters. Methods: Twenty-one cases from the publicly available LIDC-IDRI dataset were evaluated using a MONAI RetinaNet model pretrained on LUNA16 (fold 0, no fine-tuning). Five imaging conditions were tested: baseline, 25% dose reduction, 50% dose reduction, 3 mm slice thickness, and 5 mm slice thickness. Dose reduction was simulated via image-domain Gaussian noise; slice thickness via moving average along the z-axis. Detection sensitivity was computed at a confidence threshold of 0.5 with a 15 mm matching criterion. Results: Baseline sensitivity was 45.2% (57/126 consensus nodules). Dose reduction produced slight degradation: 41.3% at 25% dose and 42.1% at 50% dose. The 5 mm slice thickness condition produced a marked drop to 26.2% - a 19 percentage point reduction representing a 42% relative decrease from baseline. This finding was consistent across confidence thresholds from 0.1 to 0.9. Per-case analysis revealed heterogeneous performance including two cases with complete detection failure at baseline. Conclusion: Slice thickness represents a more fundamental constraint on AI detection performance than image noise under the conditions tested. The proposed framework is reproducible, requires no proprietary scanner data, and is designed to serve as the basis for ongoing post-deployment QA in resource-constrained environment.

AI医疗肺结节部署评估质量保证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。