arXiv:2605.30734cs.LGcs.CV2026-05

轻量模型可实现与重型模型相当的疟疾诊断性能,但解释性在实际噪声下不可靠。

Beyond Accuracy: Evaluating Efficiency, Robustness and Explainability in Deep Learning for Malaria Diagnosis

论文配图:Beyond Accuracy: Evaluating Efficiency, Robustness and Explainability in Deep Learning for Malaria Diagnosis
图 1 · 摘自论文原文
  • 对比四种不同架构模型,发现轻量模型性能不输重型模型
  • 模型准确率下降前,置信度先降低,可作为人工复核信号
  • 现有解释方法在真实临床噪声下可靠性差,影响临床可信度

疟疾仍是撒哈拉以南非洲地区的主要致死原因,当地诊断基础设施匮乏,及时准确诊断尤为困难。尽管深度学习为自动化疟疾筛查提供了可行路径,但计算成本高和决策过程不透明阻碍了临床应用。本研究在NLM-Malaria数据集上对四种涵盖广泛架构与容量的深度学习模型进行综合评估,同时考察预测性能、鲁棒性和后验可解释性。结果表明,设计高效的轻量模型在预测性能上与重型模型无显著差异(Friedman检验未达统计显著),基于CAM的XAI方法能稳定定位诊断相关区域,而细粒度归因方法在使用重型主干网络时解释更分散。在三种图像退化条件下进行鲁棒性测试显示,模型置信度下降速度超过准确率,可作为人工干预的有效信号。然而,所有XAI方法在临床可行的噪声水平下解释可靠性均显著下降,即便预测仍准确。这些发现支持在资源受限环境中部署轻量级架构,同时强调后验解释的脆弱性是负责任临床部署需关注的关键问题。

原文摘要 · Abstract (English)

Malaria remains a leading cause of mortality in sub-Saharan Africa, where scarce diagnostic infrastructure makes timely, accurate diagnosis particularly challenging. While deep learning offers a compelling path toward automated malaria screening, clinical adoption is hindered by computational cost and opacity in decision-making. This work benchmarks four deep learning models spanning a wide range of designed design architectures and model capacities on the NLM-Malaria dataset, jointly evaluating predictive performance, robustness, and post-hoc explainability. We find that lightweight, efficient-by-design models match their heavier counterparts in predictive performance, and the Friedman test confirms no statistically significant performance differences. CAM-based XAI methods consistently localize diagnostically relevant regions, while fine-grained attribution methods produce less targeted explanations, particularly with heavier backbones. Robustness evaluation under three types of image corruption further reveals that model confidence degrades faster than accuracy, providing a practical signal for human review. However, no XAI method is robust to corruption, with explanation reliability degrading at noise levels plausible in clinical practice, even when predictions remain accurate. These findings support the deployment of lightweight architectures for malaria diagnosis in resource-constrained settings, while highlighting the vulnerability of post-hoc explanations as an important consideration for responsible clinical deployment.

疟疾诊断轻量模型可解释性鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。