arXiv:2604.05515cs.CV2026-04

通过几何注意力与非空体素化,提升3D医学图像分割的精度与效率。

Geometrical Cross-Attention and Nonvoid Voxelization for Efficient 3D Medical Image Segmentation

论文配图:Geometrical Cross-Attention and Nonvoid Voxelization for Efficient 3D Medical Image Segmentation
图 1 · 摘自论文原文
  • 基于三向动态非空体素变换器,自适应聚焦关键解剖区域。
  • 在多个数据集上实现最高精度,且计算量降低56%以上。
  • 适合需要高效高精度分割的临床场景,如肿瘤与器官建模。

精准分割3D医学影像对临床诊断与治疗规划至关重要,但现有方法难以在多样解剖结构和成像模态下兼顾高精度与计算效率。为此,我们提出GCNV-Net框架,融合三向动态非空体素变压器(3DNVT)、几何交叉注意力模块(GCA)与非空体素化技术。3DNVT沿横断、矢状、冠状三个解剖平面动态划分相关体素,有效建模复杂三维空间依赖关系;GCA在多尺度特征融合中显式引入几何位置信息,显著提升细粒度解剖分割精度;非空体素化仅处理有效区域,大幅减少冗余计算,在不牺牲分割质量的前提下,相比传统体素化实现56.13%的浮点运算量(FLOPs)降低与68.49%的推理延迟下降。我们在BraTS2021、ACDC、MSD Prostate、MSD Pancreas和AMOS2022等多个基准上评估,方法在所有数据集上均达到当前最优性能,优于最佳现有方法0.65%(Dice)、0.63%(IoU)、1%(NSD),且HD95指标相对提升14.5%。结果表明GCNV-Net在精度与效率间取得良好平衡,对多种器官、疾病状态及成像模态均具强鲁棒性,具备临床部署潜力。

原文摘要 · Abstract (English)

Accurate segmentation of 3D medical scans is crucial for clinical diagnostics and treatment planning, yet existing methods often fail to achieve both high accuracy and computational efficiency across diverse anatomies and imaging modalities. To address these challenges, we propose GCNV-Net, a novel 3D medical segmentation framework that integrates a Tri-directional Dynamic Nonvoid Voxel Transformer (3DNVT), a Geometrical Cross-Attention module (GCA), and Nonvoid Voxelization. The 3DNVT dynamically partitions relevant voxels along the three orthogonal anatomical planes, namely the transverse, sagittal, and coronal planes, enabling effective modeling of complex 3D spatial dependencies. The GCA mechanism explicitly incorporates geometric positional information during multi-scale feature fusion, significantly enhancing fine-grained anatomical segmentation accuracy. Meanwhile, Nonvoid Voxelization processes only informative regions, greatly reducing redundant computation without compromising segmentation quality, and achieves a 56.13% reduction in FLOPs and a 68.49% reduction in inference latency compared to conventional voxelization. We evaluate GCNV-Net on multiple widely used benchmarks: BraTS2021, ACDC, MSD Prostate, MSD Pancreas, and AMOS2022. Our method achieves state-of-the-art segmentation performance across all datasets, outperforming the best existing methods by 0.65% on Dice, 0.63% on IoU, 1% on NSD, and relatively 14.5% on HD95. All results demonstrate that GCNV-Net effectively balances accuracy and efficiency, and its robustness across diverse organs, disease conditions, and imaging modalities highlights strong potential for clinical deployment.

3D分割医学图像高效模型注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。