arXiv:2606.19651cs.AIcs.CV2026-06被引 1

提出可同时支持临床任务与可控生成的3D脑MRI分词器。

BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation

论文配图:BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation
图 1 · 摘自论文原文
  • 用冻结的3D MAE编码器提取临床信息,专用CNN解码器重建体数据
  • 在23项任务中21项优于或持平当前最优模型,支持六变量条件生成
  • 适合需要隐私保护数据生成与疾病轨迹模拟的研究者

三维脑部MRI在神经病学与神经肿瘤学中至关重要,生成模型可扩充数据不足群体、模拟疾病进展并支持隐私保护的数据共享。现有基于潜空间扩散的方法对分词器有双重需求:编码器需保留下游任务所需的临床信息,解码器需重建解剖结构精确的体积。现有以重建为导向的分词器牺牲了临床信息保留。为此,我们提出一种全体积掩码自编码器(MAE)驱动的分词器,将编码器与解码器解耦:冻结的3D MAE编码器生成具有临床意义的嵌入,专用CNN解码器从这些嵌入的线性投影中重建体素。我们在18个公开队列(共35,309个体积)、四种模态、十类疾病和200多个采集站点上预训练编码器,并在两个场景中验证其双重用途。首先,在23项线性探测基准测试中,该编码器在21项任务上优于或持平现有最优模型(BrainIAC、BrainSegFounder、MedicalNet)。其次,基于这些临床嵌入训练的条件扩散变换器(DiT)可实现六种变量的条件生成与患者特异性纵向预测。结果表明,该方法建立了一个统一的3D脑部MRI嵌入空间,兼具下游临床任务与可控生成能力。

原文摘要 · Abstract (English)

Three-dimensional (3D) brain MRI is central to clinical neurology and neuro-oncology, where generative models could augment under-represented cohorts, simulate disease trajectories, and support privacy-preserving data sharing. Latent diffusion has been the go-to solution for modeling imaging data, but it places two competing demands on the tokenizer: encoder embeddings must retain the clinical information that downstream tasks act on, and the decoder must reconstruct anatomically faithful volumes. Existing reconstruction-driven tokenizers achieve the second at the expense of the first. To address this, we introduce a fully volumetric masked-autoencoder (MAE) based tokenizer for 3D brain MRI latent diffusion, decoupling encoder and decoder: a frozen 3D MAE encoder produces clinically informative embeddings, while a dedicated CNN decoder reconstructs voxels from a linear projection of those embeddings. We pretrain the encoder on 35,309 volumes from 18 public cohorts spanning four modalities, ten disease categories, and 200+ acquisition sites, and demonstrate its dual utility in two settings. First, on a 23-task linear-probing benchmark, the encoder outperforms or matches SOTA models (i.e., BrainIAC, BrainSegFounder, and MedicalNet) on 21 of 23 tasks. Second, a conditional diffusion transformer (DiT) trained on these clinically informative embeddings supports both conditional generation across six variables and patient-specific longitudinal forecasting. Together these results establish a single 3D brain-MRI embedding space capable of both downstream clinical tasks and controllable generation.

3D生成脑影像生成模型医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。