arXiv:2608.28787cs.CV2026-08

用预测生成框架合成3D脑部MRI,提升医学影像质量与下游任务表现。

Beyond Representation Learning: A Systematic Study of Joint-Embedding Predictive Generation for 3D Brain MRI

论文配图:Beyond Representation Learning: A Systematic Study of Joint-Embedding Predictive Generation for 3D Brain MRI
图 1 · 摘自论文原文
  • 基于3D变分自编码器生成连续隐向量,结合掩码预测与扩散机制进行生成。
  • 合成数据预训练使分类AUC提升至0.85(原0.63),分割Dice达0.80。
  • 适合需要高质量合成医学影像的研究者或医疗AI开发者。

联合嵌入预测架构(JEPAs)主要用于自监督表征学习。去噪JEPA(D-JEPA)在自然图像上展现出强大生成能力,但其在3D医学影像中的应用尚未探索。本文提出Med-D-JEPA,系统性地将联合嵌入预测生成应用于3D脑部MRI。Med-D-JEPA基于3D KL正则化对抗变分自编码器生成的连续隐向量,融合掩码上下文预测、表征对齐、逐标记扩散及迭代采样策略。在BraTS2019和OASIS-1数据集上评估了无条件与类别条件生成质量、下游分类性能,以及在BraTS2020上的初步全肿瘤分割。在不同生成设置下,Med-D-JEPA在保真度与多样性指标上优于或媲美多个强基线。相比使用真实样本训练,基于合成样本的预训练使分类AUC在BraTS2019上从0.63提升至0.85,在OASIS-1上从0.78提升至0.87;分割任务中Dice从0.74升至0.80,HD95从13.40降至9.56 mm。结果表明,联合嵌入预测生成是3D医学图像合成的有前景方向。

原文摘要 · Abstract (English)

Joint-embedding predictive architectures (JEPAs) have primarily been developed for self-supervised representation learning. Denoising JEPA (D-JEPA) recently demonstrated strong generative capabilities on natural images, yet the applicability to 3D medical imaging remains unexplored. Building on the D-JEPA framework, we present Med-D-JEPA, a systematic adaptation and evaluation of joint-embedding predictive generation for 3D brain MRI. Med-D-JEPA operates on continuous latent tokens produced by a 3D KL-regularized adversarial variational autoencoder, and combines masked context prediction, representation-level alignment, per-token diffusion, and iterative next-set-of-token sampling. We evaluate unconditional and class-conditional generation quality on BraTS2019 and OASIS-1 datasets; downstream classification utility; and preliminary whole-tumor segmentation on BraTS2020. Across different generation settings, Med-D-JEPA achieves superior or competitive performance compared to several strong baselines on fidelity and diversity metrics. Compared to training with real samples, Med-D-JEPA-based synthetic pretraining improves classification AUC from 0.63 to 0.85 on BraTS2019 and from 0.78 to 0.87 on OASIS-1. In the segmentation study, pretraining on Med-D-JEPA samples improves Dice from 0.74 to 0.80 and reduces HD95 from 13.40 to 9.56 mm. These findings establish joint-embedding predictive generation as a promising direction for 3D medical image synthesis and encourage further research in this direction.

3D生成医学影像自监督扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。