用生成模型统一医疗诊断,实现跨模态灵活推理
GenMed: A Pairwise Generative Reformulation of Medical Diagnostic Tasks

- 将医疗诊断建模为联合分布生成,通过扩散模型实现推理时条件调整
- 在仅2~4样本的少样本分割任务中表现优异,支持稀疏观测下的形状补全
- 适合需要快速适配新数据或跨模态场景的临床AI系统研发者
基于数据的医疗AI传统上将输入$X$映射到输出$Y$,通过学习函数$f$完成,但难以在真实临床环境中应对异构数据与多模态问题。本文提出一种根本性的生成范式:利用扩散模型建模联合分布$P(X,Y)$,并将推理重构为测试时的输出优化问题。通过引导生成过程匹配观测输入,该框架可在不修改架构或重新训练的情况下,实现灵活的梯度条件化,有效支持任意且此前未见的观测组合。大量实验表明,该方法在标准和跨模态医学图像分割、仅含2或4个训练样本的少样本分割、退化输入分割、从稀疏部分观测中进行形状补全,以及零样本应用方面均表现出色。为支持评估,我们构建并发布了基于MedShapeNet的大规模文本-形状数据集。结果凸显了生成联合建模作为可复用、任务无关医疗AI系统基础的巨大潜力。
原文摘要 · Abstract (English)
Data-driven medical AI is traditionally formulated as a discriminative mapping from input $X$ to output $Y$ via a learned function $f$, which does not generalize well across heterogeneous data and modalities encountered in real-world clinical settings. In this work, we propose a fundamentally different, generative paradigm. We model the joint distribution $P(X,Y)$ using diffusion models and reframe inference as a test-time output optimization problem. By guiding the generative process to match observed inputs, our framework enables flexible, gradient-based conditioning at inference time without architectural changes or retraining, effectively supporting arbitrary and previously unseen combinations of observations. Extensive experiments demonstrate strong performance across standard and cross-modality medical image segmentation, few-shot segmentation with only 2 or 4 training samples, degraded-input segmentation, shape completion from sparse and partial observations, and zero-shot application to demonstrate generality. To support these evaluations, we curated and released a large-scale text-shape dataset derived from MedShapeNet. Our results highlight the versatility of generative joint modeling as a foundation for reusable, task-agnostic medical AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。