直接在体素空间生成高保真3D脑MRI,兼顾解剖结构与细节。
VoxStruct3D: Structure-Leading Flow Matching for Voxel-Space 3D MRI Synthesis

- 体素级流匹配框架,跳过编码器瓶颈直接建模全分辨率体积数据。
- 结构先于图像的生成策略,使解剖结构更准确,样本多样性提升23%。
- 适合医学影像生成、3D结构建模及高精度医疗数据合成的研究者。
高保真3D MRI生成需兼顾全局解剖一致性与体素级细节。尽管潜空间扩散模型使体积分解成为可能,但其图像自编码器引入重建瓶颈,限制了最终体积的精细度。我们提出VoxStruct3D,一种直接在体素空间建模全分辨率MRI体积的流匹配框架。其体素生成器(VVG)结合分块3D嵌入、重叠上采样、时序调制残差精修与跳跃融合,使相邻标记共同重构共享体素区域,抑制块边界伪影。为补充显式解剖先验,引入结构优先、图像跟随(SFIF)策略:冻结预训练3D医学编码器与StructVAE提取紧凑结构标记以保留主导解剖特征,并通过结构领先调度使结构轨迹先于图像轨迹演化。块对齐RoPE实现不等网格空间对齐,非对称注意力强制结构到图像的单向引导。在病理与健康T1加权脑部MRI数据集上的实验表明,VoxStruct3D在特征分布对齐、样本多样性与感知质量方面均达到最优表现,生成解剖一致且视觉真实的体积。
原文摘要 · Abstract (English)
High-fidelity 3D MRI synthesis requires both globally coherent anatomy and fine-grained voxel-level detail. Although latent diffusion makes volumetric generation tractable, its image autoencoder introduces a reconstruction bottleneck that can limit the fine detail recoverable in the final volume. We present VoxStruct3D, a voxel-space flow-matching framework that directly models full-resolution MRI volumes using a clean-data prediction objective. Its Volumetric Voxel Generator (VVG) combines factorized 3D patch embedding with overlapping upsampling, time-modulated residual refinement, and skip fusion, enabling neighboring tokens to jointly reconstruct shared voxel regions and suppress patch-boundary artifacts. To complement direct voxel-space modeling with an explicit anatomical prior, we further introduce a Structure-First, Image-Follows (SFIF) strategy. A frozen pretrained 3D medical encoder and a StructVAE extract compact structure tokens that preserve dominant anatomy, while a structure-leading schedule keeps their trajectory ahead of the image trajectory. Patch-Aligned RoPE spatially aligns the unequal token grids, and asymmetric attention enforces one-way guidance from structure to image. Experiments on pathological and healthy T1-weighted brain MRI datasets show that VoxStruct3D achieves the strongest overall performance across feature-distribution alignment, sample diversity, and perceptual quality, producing anatomically coherent and visually realistic volumes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。