arXiv:2603.07399cs.CVeess.SP2026-03被引 1

用可解释的3D模型分类脑动脉瘤,兼顾准确率与临床可信度。

Interpretable Aneurysm Classification via 3D Concept Bottleneck Models: Integrating Morphological and Hemodynamic Clinical Features

  • 通过3D概念瓶颈模型将影像特征映射为可理解的形态与血流概念
  • 最高分类准确率达93.33%,且泛化能力稳定(差距<0.04)
  • 适合需要可解释性的医疗AI场景,如临床决策支持

我们针对深度学习在颅内动脉瘤分类中可靠评估的挑战,提出一种端到端的3D概念瓶颈框架,将高维神经影像特征映射为离散的形态学和血流动力学临床概念。采用预训练的3D ResNet-34和3D DenseNet-121提取CTA体积特征,经软瓶颈层表示人类可理解的临床概念,使用联合损失函数(诊断焦点损失+概念均方误差)优化,并通过分层五折交叉验证。结果表明,ResNet-34模型分类准确率达93.33% ± 4.5%,DenseNet-121为91.43% ± 5.8%;结合8次测试时增强(TTA)后,平均准确率为88.31%,且准确率-泛化差距小于0.04,证明高性能与可解释性可兼得。

原文摘要 · Abstract (English)

We are concerned with the challenge of reliably classifying and assessing intracranial aneurysms using deep learning without compromising clinical transparency. While traditional black-box models achieve high predictive accuracy, their lack of inherent interpretability remains a significant barrier to clinical adoption and regulatory approval. Explainability is paramount in medical modeling to ensure that AI-driven diagnoses align with established neurosurgical principles. Unlike traditional eXplainable AI (XAI) methods -- such as saliency maps, which often provide post-hoc, non-causal visual correlations -- Concept Bottleneck Models (CBMs) offer a robust alternative by constraining the model's internal logic to human-understandable clinical indices. In this article, we propose an end-to-end 3D Concept Bottleneck framework that maps high-dimensional neuroimaging features to a discrete set of morphological and hemodynamic concepts for aneurysm identification. We implemented this pipeline using a pre-trained 3D ResNet-34 backbone and a 3D DenseNet-121 to extract features from CTA volumes, which were subsequently processed through a soft bottleneck layer representing human-interpretable clinical concepts. The model was optimized using a joint-loss function to balance diagnostic focal loss and concept mean squared error (MSE), validated via stratified five-fold cross-validation. Our results demonstrate a peak task classification accuracy of 93.33% +/- 4.5% for the ResNet-34 architecture and 91.43% +/- 5.8% for the DenseNet-121 model. Furthermore, the implementation of 8-pass Test-Time Augmentation (TTA) yielded a robust mean accuracy of 88.31%, ensuring diagnostic stability during inference. By maintaining an accuracy-generalization gap of less than 0.04, this framework proves that high predictive performance can be achieved without sacrificing interpretability.

可解释AI脑动脉瘤3D医学影像概念瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。