一个模型同时生成医学影像和临床数据,无需重新训练即可灵活推理。
MetaVoxel: Joint Diffusion Modeling of Imaging and Clinical Metadata
- 用单一扩散过程联合建模影像与临床数据的分布
- 在10000+张MRI数据上实现图像生成、年龄预测等任务,性能媲美专用模型
- 支持任意输入组合的零样本推理,适合多任务医疗场景
现代深度学习方法在疾病分类、连续生物标志物估计到生成真实医学图像等任务中取得了显著成果。大多数方法仅针对特定预测方向和输入变量训练条件分布模型。我们提出MetaVoxel,一种联合扩散建模框架,通过学习覆盖所有变量的单一扩散过程,建模影像数据与临床元数据的联合分布。通过捕捉联合分布,MetaVoxel统一了传统上需独立条件模型的任务,并可在不进行任务特定重训练的情况下,使用任意输入子集实现灵活的零样本推理。基于来自九个数据集的超过10,000例T1加权MRI扫描及其临床元数据,我们证明单个MetaVoxel模型可完成图像生成、年龄估计和性别预测,性能与现有专用基准相当。额外实验展示了其灵活推理能力。这些结果表明,联合多模态扩散为统一医疗AI模型并提升临床适用性提供了有前景的方向。
原文摘要 · Abstract (English)
Modern deep learning methods have achieved impressive results across tasks from disease classification, estimating continuous biomarkers, to generating realistic medical images. Most of these approaches are trained to model conditional distributions defined by a specific predictive direction with a specific set of input variables. We introduce MetaVoxel, a generative joint diffusion modeling framework that models the joint distribution over imaging data and clinical metadata by learning a single diffusion process spanning all variables. By capturing the joint distribution, MetaVoxel unifies tasks that traditionally require separate conditional models and supports flexible zero-shot inference using arbitrary subsets of inputs without task-specific retraining. Using more than 10,000 T1-weighted MRI scans paired with clinical metadata from nine datasets, we show that a single MetaVoxel model can perform image generation, age estimation, and sex prediction, achieving performance comparable to established task-specific baselines. Additional experiments highlight its capabilities for flexible inference. Together, these findings demonstrate that joint multimodal diffusion offers a promising direction for unifying medical AI models and enabling broader clinical applicability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。