arXiv:2509.25265eess.IVcs.LG2025-09中稿 · ARRS 2026 Annual M…

研究噪声对胸部X光影像分割和分类模型的影响,发现分割模型极脆弱。

Evaluating the Impact of Radiographic Noise on Chest X-ray Semantic Segmentation and Disease Classification Using a Scalable Noise Injection Framework

  • 构建可扩展噪声注入框架,模拟量子与电子噪声影响
  • 严重电子噪声使肺部分割性能下降84.3%(Dice系数)
  • 分类任务对不同噪声类型有差异性脆弱,需针对性验证

深度学习模型在放射影像分析中应用日益广泛,但其可靠性受临床成像固有随机噪声的挑战。目前缺乏对不同噪声类型如何系统影响模型的跨任务理解。本文利用新型可扩展噪声注入框架,评估了先进卷积神经网络(CNN)在两种关键胸部X光任务中的鲁棒性:语义分割与肺部疾病分类。针对常见架构(UNet、DeepLabV3、FPN;ResNet、DenseNet、EfficientNet)在公开数据集(Landmark、ChestX-ray14)上施加受控且临床相关的噪声强度,包括量子(泊松)和电子(高斯)噪声。结果显示任务鲁棒性存在显著差异:语义分割模型极为脆弱,严重电子噪声下肺部分割性能崩溃(Dice相似系数下降0.843),几乎完全失效;而分类任务整体更具韧性,但非均匀——某些任务如气胸与肺不张区分在量子噪声下惨败(AUROC下降0.355),其他则更易受电子噪声影响。结果表明,尽管分类模型具备一定内在鲁棒性,像素级分割任务仍极度脆弱。模型失败具有任务与噪声特异性,强调临床部署前必须采取针对性验证与缓解策略。

原文摘要 · Abstract (English)

Deep learning models are increasingly used for radiographic analysis, but their reliability is challenged by the stochastic noise inherent in clinical imaging. A systematic, cross-task understanding of how different noise types impact these models is lacking. Here, we evaluate the robustness of state-of-the-art convolutional neural networks (CNNs) to simulated quantum (Poisson) and electronic (Gaussian) noise in two key chest X-ray tasks: semantic segmentation and pulmonary disease classification. Using a novel, scalable noise injection framework, we applied controlled, clinically-motivated noise severities to common architectures (UNet, DeepLabV3, FPN; ResNet, DenseNet, EfficientNet) on public datasets (Landmark, ChestX-ray14). Our results reveal a stark dichotomy in task robustness. Semantic segmentation models proved highly vulnerable, with lung segmentation performance collapsing under severe electronic noise (Dice Similarity Coefficient drop of 0.843), signifying a near-total model failure. In contrast, classification tasks demonstrated greater overall resilience, but this robustness was not uniform. We discovered a differential vulnerability: certain tasks, such as distinguishing Pneumothorax from Atelectasis, failed catastrophically under quantum noise (AUROC drop of 0.355), while others were more susceptible to electronic noise. These findings demonstrate that while classification models possess a degree of inherent robustness, pixel-level segmentation tasks are far more brittle. The task- and noise-specific nature of model failure underscores the critical need for targeted validation and mitigation strategies before the safe clinical deployment of diagnostic AI.

医学影像噪声鲁棒性分割模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。