通过可控制的2D切片导航预训练,提升MRI图像的空间与解剖表征能力
3D MRI Image Pretraining via Controllable 2D Slice Navigation Task

- 将3D MRI转化为可控的2D切片序列,构建动作条件的自监督任务
- 在多个解剖与空间下游任务中优于静态体积基线和无对齐动作的动态模型
- 适合大规模无标签MRI数据的表征学习,尤其对空间结构理解有帮助
自监督预训练已成为从无标签MRI扫描中学习表征的主流方法。然而,现有方法大多将每张扫描视为切片、补丁或体积分块的静态集合。本文提出一种新思路:将3D体积分解为可控制的2D渲染序列——通过连续位置、方向和尺度的切片渲染,将3D体积转化为密集视频-动作序列,其控制变量为动作轨迹。我们采用动作条件的预训练目标,由编码器对切片观测进行编码,潜态动力学模型预测潜变量演化。在代表性解剖与空间下游任务中,该方法相较于标准静态体积基线、仅使用编码器的预训练以及无对齐动作的动力学变体均表现更优。结果表明,可控的MRI切片导航为从大规模无标签MRI数据中学习解剖与空间表征提供了有效补充接口。
原文摘要 · Abstract (English)
Self-supervised pretraining has become the mainstream approach for learning MRI representations from unlabeled scans. However, most existing objectives still treat each scan primarily as static aggregations of slices, patches or volumes. We ask whether there exists an intrinsic form of self-supervision signal that is different from reconstructing the masked patches, through transforming the 3D volumes into controllable 2D rendered sequences: by rendering slices at continuous positions, orientations, and scales, a 3D volume can be converted into dense video-action sequences whose controls are the action trajectories. We study this formulation with an action-conditioned pretraining objective, where a tokenizer encodes slice observations and a latent dynamics model predicts the evolution of latent features. Across representative anatomical and spatial downstream tasks, the proposed pretraining is evaluated against standard static-volume baselines, tokenizer-only pretraining, and dynamics variants without aligned actions. These results suggest that controllable MRI slice navigation provides a useful complementary pretraining interface for learning anatomical and spatial representations from large unlabeled MRI collections.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。