用少量步骤实现高效精准的医学图像分割。
RF-HiT: Rectified Flow Hierarchical Transformer for General Medical Image Segmentation

- 基于修正流与分层变换器,降低计算复杂度。
- 仅需3步推理,参数量1360万,性能达91.27% Dice。
- 适合临床部署,兼顾速度与精度。
准确的医学图像分割需要长程上下文推理和精确边界识别,而现有基于Transformer或扩散模型的方法常受限于二次计算复杂度和高昂推理延迟。我们提出RF-HiT,一种结合倒置金字塔变换器主干与多尺度分层编码器的修正流分层Transformer,实现解剖结构引导的特征调节。与依赖数百次去噪步骤的扩散方法不同,RF-HiT采用修正流与高效变换器模块,实现线性复杂度,仅需少数离散化步骤。模型通过可学习插值在各分辨率融合条件特征,实现高效的多尺度特征整合且计算开销极小。结果表明,RF-HiT仅需10.14 GFLOPs、1360万参数,推理仅3步即可达成91.27% mean Dice(ACDC)与87.40%(BraTS 2021),性能媲美甚至超越更复杂的架构。这表明RF-HiT是临床图像分割极具潜力的高效基础模型。
原文摘要 · Abstract (English)
Accurate medical image segmentation requires both long-range contextual reasoning and precise boundary delineation, a task where existing transformer- and diffusion-based paradigms are frequently bottlenecked by quadratic computational complexity and prohibitive inference latency. We propose RF-HiT, a Rectified Flow Hierarchical Transformer that integrates an Hourglass Transformer backbone with a multi-scale hierarchical encoder for anatomically guided feature conditioning. Unlike prior diffusion-based approaches that rely on hundreds of denoising steps, RF-HiT leverages rectified flow with efficient transformer blocks, achieving linear complexity and requiring only a few discretization steps. The model further fuses conditioning features at each resolution via learnable interpolation, enabling effective multi-scale feature integration with minimal computational overhead. As a result, RF-HiT achieves a strong efficiency-performance trade-off, requiring only 10.14 GFLOPs, 13.6M parameters, and inference in as few as 3 steps. Despite its compact design, RF-HiT attains 91.27% mean Dice on ACDC and 87.40% on BraTS 2021, achieving performance comparable to or exceeding that of significantly more intensive architectures. These results suggest that RF-HiT is a promising, computationally efficient foundation for clinical image segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。