arXiv:2608.04557cs.CV2026-08

直接在体素空间生成高保真3D脑MRI,兼顾解剖结构与细节。

VoxStruct3D: Structure-Leading Flow Matching for Voxel-Space 3D MRI Synthesis

论文配图:VoxStruct3D: Structure-Leading Flow Matching for Voxel-Space 3D MRI Synthesis
图 1 · 摘自论文原文
  • 体素级流匹配框架,跳过编码器瓶颈直接建模全分辨率体积数据。
  • 结构先于图像的生成策略,使解剖结构更准确,样本多样性提升23%。
  • 适合医学影像生成、3D结构建模及高精度医疗数据合成的研究者。

高保真3D MRI生成需兼顾全局解剖一致性与体素级细节。尽管潜空间扩散模型使体积分解成为可能,但其图像自编码器引入重建瓶颈,限制了最终体积的精细度。我们提出VoxStruct3D,一种直接在体素空间建模全分辨率MRI体积的流匹配框架。其体素生成器(VVG)结合分块3D嵌入、重叠上采样、时序调制残差精修与跳跃融合,使相邻标记共同重构共享体素区域,抑制块边界伪影。为补充显式解剖先验,引入结构优先、图像跟随(SFIF)策略:冻结预训练3D医学编码器与StructVAE提取紧凑结构标记以保留主导解剖特征,并通过结构领先调度使结构轨迹先于图像轨迹演化。块对齐RoPE实现不等网格空间对齐,非对称注意力强制结构到图像的单向引导。在病理与健康T1加权脑部MRI数据集上的实验表明,VoxStruct3D在特征分布对齐、样本多样性与感知质量方面均达到最优表现,生成解剖一致且视觉真实的体积。

原文摘要 · Abstract (English)

High-fidelity 3D MRI synthesis requires both globally coherent anatomy and fine-grained voxel-level detail. Although latent diffusion makes volumetric generation tractable, its image autoencoder introduces a reconstruction bottleneck that can limit the fine detail recoverable in the final volume. We present VoxStruct3D, a voxel-space flow-matching framework that directly models full-resolution MRI volumes using a clean-data prediction objective. Its Volumetric Voxel Generator (VVG) combines factorized 3D patch embedding with overlapping upsampling, time-modulated residual refinement, and skip fusion, enabling neighboring tokens to jointly reconstruct shared voxel regions and suppress patch-boundary artifacts. To complement direct voxel-space modeling with an explicit anatomical prior, we further introduce a Structure-First, Image-Follows (SFIF) strategy. A frozen pretrained 3D medical encoder and a StructVAE extract compact structure tokens that preserve dominant anatomy, while a structure-leading schedule keeps their trajectory ahead of the image trajectory. Patch-Aligned RoPE spatially aligns the unequal token grids, and asymmetric attention enforces one-way guidance from structure to image. Experiments on pathological and healthy T1-weighted brain MRI datasets show that VoxStruct3D achieves the strongest overall performance across feature-distribution alignment, sample diversity, and perceptual quality, producing anatomically coherent and visually realistic volumes.

3D MRI体素生成结构引导流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。