arXiv:2409.08691cs.CV2024-09AAAI被引 5

用自回归建模3D医学图像,提升上下文理解能力

Autoregressive Sequence Modeling for 3D Medical Image Representation

  • 将3D医学图像按空间/对比/语义关系序列化为视觉标记
  • 在9个下游任务中优于现有方法,提升表示学习效果
  • 适合需要精细上下文理解的医学图像分析场景

三维(3D)医学影像如计算机断层扫描(CT)和磁共振成像(MRI)在临床中至关重要。然而,由于器官、诊断任务和成像模态的多样性,对多样化且全面的表征需求尤为突出。如何有效解析复杂的上下文信息并从中提取有意义的洞察仍是该领域的开放挑战。尽管当前自监督学习方法展现出潜力,但通常将图像视为整体,忽略了单幅或多幅图像中局部区域间的复杂关系。本文提出一种开创性的自回归预训练框架,用于学习3D医学图像表征。该方法基于空间、对比度和语义相关性对多种3D医学图像进行序列化,将其视为序列中的互联视觉标记。通过自回归序列建模任务预测下一个视觉标记,使模型深入理解并整合3D医学图像中的上下文信息。此外,我们引入随机起始策略以避免高估标记间关系,增强学习鲁棒性。在公开数据集上的九个下游任务中,该方法表现出优于现有方法的性能。

原文摘要 · Abstract (English)

Three-dimensional (3D) medical images, such as Computed Tomography (CT) and Magnetic Resonance Imaging (MRI), are essential for clinical applications. However, the need for diverse and comprehensive representations is particularly pronounced when considering the variability across different organs, diagnostic tasks, and imaging modalities. How to effectively interpret the intricate contextual information and extract meaningful insights from these images remains an open challenge to the community. While current self-supervised learning methods have shown potential, they often consider an image as a whole thereby overlooking the extensive, complex relationships among local regions from one or multiple images. In this work, we introduce a pioneering method for learning 3D medical image representations through an autoregressive pre-training framework. Our approach sequences various 3D medical images based on spatial, contrast, and semantic correlations, treating them as interconnected visual tokens within a token sequence. By employing an autoregressive sequence modeling task, we predict the next visual token in the sequence, which allows our model to deeply understand and integrate the contextual information inherent in 3D medical images. Additionally, we implement a random startup strategy to avoid overestimating token relationships and to enhance the robustness of learning. The effectiveness of our approach is demonstrated by the superior performance over others on nine downstream tasks in public datasets.

3D医学图像自回归建模表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。