用多层反向扫描与频域协同注意力,提升小血管分割精度
Polygon-mamba: Retinal vessel segmentation using polygon scanning mamba and space-frequency collaborative attention
- 设计多层反向扫描Mamba模块,保持小血管拓扑连续性
- 在三个数据集上F1超0.82,AUC超0.98,敏感度达0.84
- 适合眼科影像分析、微血管病灶检测等临床应用
视网膜血管分割对眼病诊断至关重要,尤其小血管分割仍具挑战。为此,我们提出融合CNN与Mamba的混合网络,引入多层反向扫描视觉状态空间模型(PS-VSS)以识别小血管结构特征,有效保留像素连通性,减少信息丢失。同时,在跳跃连接中设计空间-频率协同注意力机制(SFCAM),联合提取空间域的位置结构信息与频率域的全局感知及局部细节,动态增强关键特征并抑制噪声。模型在DRIVE、STARE和CHASE_DB1三个公开数据集上测试,相较于人工标注,分别取得F1分数0.8283、0.8282、0.8251,AUC值0.9806、0.9840、0.9866,敏感度(SE)为0.8268、0.8314、0.8484。结果通过可视化与定量分析验证了模型有效性。
原文摘要 · Abstract (English)
Retinal vessel segmentation is crucial for diagnosis and assessment of ocular diseases. Notably, segmentation of small retinal vessels has been consistently recognized as a challenging and complex task. To tackle this challenge, we design a hybrid CNN-Mamba fusion network that integrates polygon scanning mamba and space-frequency collaborative attention mechanism for the detection of small vessels. Considering that the traditional mamba architecture with horizontal-vertical scanning may compromise the topological integrity of target structures and result in local discontinuities in small retinal vessels, we present a polygon scanning visual state space model (PS-VSS) to identify small vessel structural features by multi-layer reverse scanning way. Which effectively preserves pixels connectivity, thereby substantially mitigating the loss of information pertaining to small vessels. Furthermore, as we all known that the spatial domain prioritizes positional and structural information, while the frequency domain emphasizes global perception and local detail components, a space-frequency collaborative attention mechanism (SFCAM) is introduced within the skip connection to extract efficient features from the spatial and frequency domains. This strategy empowers the model to dynamically enhance the key features while effectively suppressing clutters. To assess the efficacy of our model, it was tested on three publicly available datasets: DRIVE, STARE, and CHASE_DB1. Compared to manual annotations, our model demonstrated F1 scores of 0.8283, 0.8282, and 0.8251, Area Under Curve (AUC) values of 0.9806, 0.9840, and 0.9866, and Sensitivity (SE) values of of 0.8268, 0.8314, and 0.8484 across three datasets, respectively. The effectiveness of our model was validated through both visual inspection and quantitative analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。