用带几何信息的高斯令牌提升脑肿瘤3D影像压缩效率
GSToken: Geometry-Structured Gaussian Tokens for Compact 3D Medical Image Representation

- 用高斯分布参数建模每个令牌的中心、尺度和方向
- 在固定探针下,性能显著优于传统自适应方法
- 适合需要紧凑且几何精确的医学影像分析场景
多模态MRI的有效分割对提升脑肿瘤识别神经网络精度至关重要。现有方法通常通过固定块编码或学习注意力池化(如TokenLearner)将3D体数据压缩为令牌序列,但这类方案丢弃了显式空间形状信息,导致令牌无法表达病灶形态与空间范围。同时,端到端评估将分词器的信息保留能力与下游解码器重建能力混杂,缺乏统一容量协议使性能差异难以归因。本文首次在多模态脑肿瘤分割中引入高斯令牌:每个令牌不仅携带语义特征,还包含可学习的3D中心、各向异性尺度与方向,以极低参数开销赋予表示显式几何支持。我们提出冻结令牌评估协议:训练后冻结分词器,其输出转化为固定容量序列化合约,使用共享轻量Transformer探针在严格匹配条件下独立衡量各分词器保留的信息量。多种子配对统计检验表明,GSToken在冻结探针下始终显著优于容量匹配的自适应基线,且在所有肿瘤亚区域、表面及距离度量上均具一致性优势。结果证明,在令牌中显式编码空间几何能显著提升体数据表示的信息密度,为紧凑3D医学图像表示提供了新设计范式。
原文摘要 · Abstract (English)
Effective segmentation of multi-modal MRI is central to improving neural network accuracy in brain tumor recognition. Existing methods typically compress 3D volumes into token sequences via fixed patch encoding or learned attention pooling (e.g., TokenLearner). However, these compression schemes discard explicit spatial shape information; the resulting tokens convey no notion of lesion morphology or spatial extent. Meanwhile, end-to-end evaluation entangles a tokenizer's information retention with the reconstruction capacity of the downstream decoder, and the lack of a unified capacity contract across methods makes performance differences difficult to attribute. In this paper, we introduce Gaussian tokens to multi-modal brain tumor segmentation for the first time: each token carries not only a semantic feature but also a learned 3D center, anisotropic scale, and orientation, endowing the representation with explicit geometric support at negligible parameter cost. We further propose a frozen-token utility evaluation protocol: the trained tokenizer is frozen, its output is cast into a fixed-capacity serialized contract, and a shared lightweight Transformer probe independently measures each tokenizer's retained information under strictly matched conditions. Multi-seed paired statistical testing shows that GSToken consistently and substantially outperforms capacity-matched adaptive baselines under frozen probing, with uniform advantages across all tumor sub-regions, surface, and distance metrics. These results demonstrate that explicitly encoding spatial geometry within tokens significantly improves the information density of volumetric representations, offering a new design principle for compact 3D medical image representation and downstream reading.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。