融合拓扑分析与视觉变换器,实现99%以上准确率的脑肿瘤分类。
Bridging Topology and Deep Representation Learning: A TDA-ViT Fusion Model for Four-Class Brain Tumor Classification

- 用拓扑数据分析提取肿瘤结构特征,与视觉变换器的语义特征融合。
- 在BRISC2025数据集上达99.10%准确率,各项指标均超现有模型。
- 适合医学影像智能诊断研究者,尤其关注结构特征建模的场景。
从磁共振成像(MRI)中精确分类脑肿瘤是早期诊断和临床决策的关键。视觉变换器(ViTs)在学习全局上下文表征方面表现优异,但难以捕捉肿瘤区域固有的结构与拓扑模式。为此,我们提出一种融合框架,将拓扑数据分析(TDA)特征与预训练视觉变换器表示相结合,用于四类脑肿瘤分类。该方法利用TDA提取图像中的几何结构、连通性与形状等互补拓扑描述符;同时,预训练ViT模型从相同图像中学习高层语义表示。两者特征空间融合后形成统一且更具判别性的表征。在包含胶质瘤、脑膜瘤、垂体瘤及非肿瘤四类样本的BRISC2025数据集上评估,结果表明:融合模型性能显著优于单一方法,准确率达99.10%,精确率为99.27%,召回率为99.15%,F1得分为99.21%,AUC为99.98%。该模型超越了ResNet50、ResNet101、EfficientNetB2及独立视觉变换器等先进模型,证明拓扑特征可为深度表征学习提供关键补充信息,构建出鲁棒高效的自动化脑肿瘤分类框架。
原文摘要 · Abstract (English)
Accurate brain tumor classification from magnetic resonance imaging (MRI) is a key requirement for early diagnosis and clinical decision-making. Vision Transformers (ViTs) have shown strong performance in medical image analysis by learning global contextual representations, but they often fail to capture intrinsic structural and topological patterns present in tumor regions. To address this limitation, we propose a fusion framework that combines Topological Data Analysis (TDA) features with pretrained Vision Transformer representations for four-class brain tumor classification. In the proposed method, TDA is used to extract complementary topological descriptors that capture geometric structure, connectivity, and shape information from MRI images. In parallel, a pretrained ViT model learns high-level semantic representations from the same images. These two feature spaces are then fused to form a unified and more discriminative representation for classification. The model is evaluated on the BRISC2025 dataset, which contains four brain tumor classes: glioma, meningioma, pituitary tumor, and non-tumor cases. Experimental results show that combining topological and transformer-based features significantly improves performance compared to using either approach alone. The proposed TDA-ViT fusion model achieves an accuracy of 99.10%, precision of 99.27%, recall of 99.15%, F1-score of 99.21%, and an AUC of 99.98%. It also outperforms several state-of-the-art models, including ResNet50, ResNet101, EfficientNetB2, and standalone Vision Transformers. These results demonstrate that topological features provide valuable complementary information that enhances deep representation learning, leading to a robust and highly accurate framework for automated brain tumor classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。