融合空间与频域特征,提升医学图像分类准确率
S$^3$F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network
- 双分支架构同时学习图像的空间与频域特征
- 在多个数据集上最高提升5.13%准确率,达98.76%
- 可动态调整依赖,适合病理复杂场景的分析
卷积神经网络在医学图像分析中因擅长提取层次化空间特征而成为主流,但其单一空间视角难以捕捉全局模式,也未显式建模频域特性。为此,我们提出空间-频谱摘要融合网络(S$^3$F-Net),一种双分支框架,同时从空间和频域表征中学习。S$^3$F-Net将深层空间CNN与我们提出的浅层频谱编码器SpectraNet融合。SpectraNet引入频谱滤波层(SpectralFilter),基于卷积定理通过高效逐元素乘法直接作用于图像完整傅里叶谱,实现瞬时全局感受野,其输出由轻量级摘要网络提炼。我们在四个跨模态医学图像数据集上评估该框架,结果表明其在所有情况下均显著优于仅用空间特征的基线模型,准确率最高提升5.13%。采用双线性融合的S$^3$F-Net在BRISC2025数据集上达到98.76%的SOTA级准确率;在纹理主导的胸部X光肺炎数据集上,拼接融合策略表现更佳,达93.11%准确率,超越许多更深模型。可解释性分析显示,S$^3$F-Net能根据输入病理动态调节各分支权重。这些结果验证了双域方法在医学图像分析中的强大通用性。
原文摘要 · Abstract (English)
Convolutional Neural Networks have become a cornerstone of medical image analysis due to their proficiency in learning hierarchical spatial features. However, this focus on a single domain is inefficient at capturing global, holistic patterns and fails to explicitly model an image's frequency-domain characteristics. To address these challenges, we propose the Spatial-Spectral Summarizer Fusion Network (S$^3$F-Net), a dual-branch framework that learns from both spatial and spectral representations simultaneously. The S$^3$F-Net performs a fusion of a deep spatial CNN with our proposed shallow spectral encoder, SpectraNet. SpectraNet features the proposed SpectralFilter layer, which leverages the Convolution Theorem by applying a bank of learnable filters directly to an image's full Fourier spectrum via a computation-efficient element-wise multiplication. This allows the SpectralFilter layer to attain a global receptive field instantaneously, with its output being distilled by a lightweight summarizer network. We evaluate S$^3$F-Net across four medical imaging datasets spanning different modalities to validate its efficacy and generalizability. Our framework consistently and significantly outperforms its strong spatial-only baseline in all cases, with accuracy improvements of up to 5.13%. With a powerful Bilinear Fusion, S$^3$F-Net achieves a SOTA competitive accuracy of 98.76% on the BRISC2025 dataset. Concatenation Fusion performs better on the texture-dominant Chest X-Ray Pneumonia dataset, achieving 93.11% accuracy, surpassing many top-performing, much deeper models. Our explainability analysis also reveals that the S$^3$F-Net learns to dynamically adjust its reliance on each branch based on the input pathology. These results verify that our dual-domain approach is a powerful and generalizable paradigm for medical image analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。