提出用MLP替代传统模型,更有效捕捉医学图像的细粒度长程依赖。
Toward Next-generation Medical Vision Backbones: Modeling Finer-grained Long-range Visual Dependency
- 用MLP建模高分辨率医学图像中的细粒度长程依赖
- 在多种任务上超越CNN与Transformer,提升性能
- 为下一代医学视觉骨干网络提供新范式
医学图像计算(MIC)涵盖像素级(如分割、配准)和图像级(如分类、回归)任务。有效分析需同时捕捉全局长程上下文与局部细微特征,因此需精细建模长程视觉依赖。相比受固有局部性限制的卷积神经网络(CNN),Transformer擅长长程建模,但自注意力计算开销大,通常无法处理高分辨率特征(如未下采样前的全图特征或补丁嵌入),难以建模医学图像中细微结构的长程依赖。与此同时,基于多层感知机(MLP)的视觉模型被证明在计算与内存效率上更具优势,但在MIC领域尚未广泛研究。本博士研究推进了深度学习在医学图像计算中的应用,首次将Transformer应用于像素级与图像级任务;随后聚焦于MLP,开创性地构建了基于MLP的视觉模型,以捕捉医学图像中细粒度的长程依赖。大量实验验证了长程依赖建模在MIC中的关键作用,并揭示一个重要发现:MLP可在保留丰富解剖/病理细节的高分辨率特征中,实现对细粒度长程依赖的有效建模。该发现确立了MLP优于传统Transformer与CNN的范式,在多种医学视觉任务中持续提升性能,为下一代医学视觉骨干网络铺平道路。
原文摘要 · Abstract (English)
Medical Image Computing (MIC) is a broad research topic covering both pixel-wise (e.g., segmentation, registration) and image-wise (e.g., classification, regression) vision tasks. Effective analysis demands models that capture both global long-range context and local subtle visual characteristics, necessitating fine-grained long-range visual dependency modeling. Compared to Convolutional Neural Networks (CNNs) that are limited by intrinsic locality, transformers excel at long-range modeling; however, due to the high computational loads of self-attention, transformers typically cannot process high-resolution features (e.g., full-scale image features before downsampling or patch embedding) and thus face difficulties in modeling fine-grained dependency among subtle medical image details. Concurrently, Multi-layer Perceptron (MLP)-based visual models are recognized as computation/memory-efficient alternatives in modeling long-range visual dependency but have yet to be widely investigated in the MIC community. This doctoral research advances deep learning-based MIC by investigating effective long-range visual dependency modeling. It first presents innovative use of transformers for both pixel- and image-wise medical vision tasks. The focus then shifts to MLPs, pioneeringly developing MLP-based visual models to capture fine-grained long-range visual dependency in medical images. Extensive experiments confirm the critical role of long-range dependency modeling in MIC and reveal a key finding: MLPs provide feasibility in modeling finer-grained long-range dependency among higher-resolution medical features containing enriched anatomical/pathological details. This finding establishes MLPs as a superior paradigm over transformers/CNNs, consistently enhancing performance across various medical vision tasks and paving the way for next-generation medical vision backbones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。