量化医学影像模型性能下降的因果机制,帮医生选对优化方向。
Causal Attribution of Model Performance Gaps in Medical Imaging Under Distribution Shifts
- 用因果图+谢林值分解成像协议和标注差异的影响
- 跨中心时采集差异导致性能降6.5%±9.1%,跨标注者时标注差异占7.4%±8.9%
- 适合部署前分析模型失败原因,指导针对性改进
医学影像分割的深度学习模型在分布偏移下性能显著下降,但其因果机制尚不明确。本文将因果归因框架扩展至高维分割任务,量化成像协议与标注变异性对性能退化的独立贡献。通过构建数据生成的因果图,并使用谢林值公平分配性能变化,解决医学影像中的高维输出、样本有限与机制复杂交互等挑战。在多发性硬化症(MS)病灶分割任务中,跨4个中心与7名标注者验证发现:跨标注者时,标注协议变动主导(7.4% ± 8.9% DSC归因),跨中心时则为采集差异主导(6.5% ± 9.1%)。该机制级量化可帮助实践者根据部署场景优先采取针对性干预。
原文摘要 · Abstract (English)
Deep learning models for medical image segmentation suffer significant performance drops due to distribution shifts, but the causal mechanisms behind these drops remain poorly understood. We extend causal attribution frameworks to high-dimensional segmentation tasks, quantifying how acquisition protocols and annotation variability independently contribute to performance degradation. We model the data-generating process through a causal graph and employ Shapley values to fairly attribute performance changes to individual mechanisms. Our framework addresses unique challenges in medical imaging: high-dimensional outputs, limited samples, and complex mechanism interactions. Validation on multiple sclerosis (MS) lesion segmentation across 4 centers and 7 annotators reveals context-dependent failure modes: annotation protocol shifts dominate when crossing annotators (7.4% $\pm$ 8.9% DSC attribution), while acquisition shifts dominate when crossing imaging centers (6.5% $\pm$ 9.1%). This mechanism-specific quantification enables practitioners to prioritize targeted interventions based on deployment context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。