提出结构感知预训练与频域边界优化结合的新框架,提升医学图像分割在小样本下的精度与边界清晰度。
SpectraFlow: Unifying Structural Pretraining and Frequency Adaptation for Medical Image Segmentation

- 分两阶段:先用混合域流形对齐学习结构特征,再通过注意力融合与频向动态卷积优化边界。
- 在ISIC-2016等3个数据集上实现优于现有方法的分割精度,小样本下边界更清晰。
- 适合医疗影像小样本分割场景,尤其关注边界细节和几何一致性研究者。
医学图像分割在低数据条件下仍具挑战性,稀疏标注常导致泛化能力差、边界模糊及细结构缺失。近期自监督预训练虽提升了迁移性,但易产生纹理偏差。准确分割本质上依赖几何感知,需兼顾拓扑一致性和边界精确性。为此,我们提出两阶段框架:第一阶段通过混合域均值流形预训练(Mixed-Domain MeanFlow Pretraining),在共享潜在空间中对齐图像与二值掩码,以掩码作为结构引导而非预测目标,实现任务无关预训练;为增强稀缺监督下的训练稳定性,引入轻量级分散损失防止表示坍缩。第二阶段使用轻量解码器,结合直接注意力融合实现跨尺度自适应门控,以及频向动态卷积,在外观变化下强化高频边界细节。在ISIC-2016、Kvasir-SEG和GlaS上的实验表明,该方法持续优于现有先进方法,尤其在低数据设置下具备更强鲁棒性与更锐利的边界分割能力。
原文摘要 · Abstract (English)
Medical image segmentation remains challenging in low-data regimes, where scarce annotations often yield poor generalization and ambiguous boundaries with missing fine structures. Recent self-supervised pretraining has improved transferability, but it often exhibits a texture bias. In contrast, accurate segmentation is inherently geometry-aware and depends on both topological consistency and precise boundary preservation. To address this problem, we propose a two-stage framework that couples structure-aware encoder pretraining with boundary-oriented decoding. In Stage-1, we aim to learn structure-aware representations for downstream segmentation in low-data regimes. To this end, we propose Mixed-Domain MeanFlow Pretraining, which aligns images and binary masks in a shared latent space through latent transport regression, where masks act as conditional structural guidance rather than prediction targets, making the pretraining task-agnostic. To further improve training stability under scarce supervision, we incorporate a lightweight Dispersive Loss to prevent representation collapse. In Stage-2, we fine-tune the pretrained encoder with a lightweight decoder that combines Direct Attentional Fusion for adaptive cross-scale gating and Frequency-Directional Dynamic Convolution for high-frequency boundary refinement under appearance variation. Experiments on ISIC-2016, Kvasir-SEG, and GlaS demonstrate consistent gains over state-of-the-art methods, with improved robustness in low-data settings and sharper boundary delineation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。