用因果生成模型模拟真实临床变化,更准确评估医学影像模型的鲁棒性。
Counterfactual Stress Testing for Image Classification Models

- 基于因果生成模型生成'如果'场景图像,控制变量干预扫描设备和性别等属性
- 在胸片和乳腺钼靶中验证,性能下降方向与真实分布偏移一致,排名相关性更强
- 适合医疗AI部署前评估,尤其关注数据分布偏移的模型开发者
医学影像中的深度学习模型在新临床环境中常因人口统计、扫描设备或采集协议的分布偏移而失效。核心挑战在于模型的不充分指定问题:验证性能相近的模型在真实世界中表现出截然不同的失败模式。尽管应力测试已成评估手段,但现有方法多依赖简单、无指导的扰动(如亮度或对比度变化),无法捕捉临床真实变异,易高估鲁棒性。本文提出一种基于因果生成模型的反事实应力测试框架,通过干预扫描类型、记录性别等属性,在基本保持解剖身份的前提下生成真实的‘如果’图像,实现针对特定分布偏移的受控且语义明确的评估。在两种成像模态(胸部X光、乳腺钼靶)、三种模型架构及多个偏移场景下,结果表明反事实应力测试比传统扰动能更准确地预测真实分布外性能,捕捉了性能变化的方向与相对幅度,并展现出更强的整体排名一致性。这表明因果生成模型可为部署前评估目标分布偏移下的模型鲁棒性提供信息丰富的合成压力测试。
原文摘要 · Abstract (English)
Deep learning models in medical imaging often fail when deployed in new clinical environments due to distribution shifts in demographics, scanner hardware, or acquisition protocols. A central challenge is underspecification, where models with similar validation performance exhibit divergent real-world failure modes. Although stress testing has emerged as a tool to assess this, current methods typically rely on simple, uninformed perturbations (e.g., brightness or contrast changes), which fail to capture clinically realistic variation and can overestimate robustness. In this work, we introduce a counterfactual stress testing framework based on causal generative models that create realistic "what if" images by intervening on attributes such as scanner type and recorded sex while largely preserving anatomical identity, enabling controlled and semantically meaningful evaluation under targeted distribution shifts. Across two imaging modalities (chest X-ray and mammography), three model architectures, and multiple shift scenarios, we show that counterfactual stress tests provide a substantially more accurate proxy for real out-of-distribution performance than classical perturbations, capturing the direction and relative magnitude of performance changes and showing stronger overall rank agreement. These results suggest that causal generative models can provide informative synthetic stress tests for assessing robustness under targeted distribution shifts prior to deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。