arXiv:2608.08135cs.CVcs.AI2026-08中稿 · ad Sashimi 2026

用统一模型实现全体积跨模态医学图像生成,支持零样本泛化。

Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching

论文配图:Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching
图 1 · 摘自论文原文
  • 基于预训练3D变分自编码器构建体积先验,将翻译转为条件流匹配问题。
  • 单模型在三组数据上完成跨模态与同模态任务,性能媲美专用模型。
  • 可零样本迁移至未见解剖区域,支持组合式跨数据集图像合成。

跨模态医学图像生成可减轻多模态扫描负担,但受限于两大耦合问题:现有方法仅处理2D切片或3D小块而非完整体积,且每个任务需独立建模。根源在于缺乏足够强的体积先验,导致生成模型需同时学习解剖外观与跨模态映射,而可用配对数据量不足,使该问题在规模上不成立。本文提出解耦目标:利用大规模预训练3D变分自编码器提供紧凑的体积外观隐表示,将翻译简化为条件流匹配问题。此压缩使全体积处理可行,分辨率感知采样策略保留原生解剖尺度。我们在三个多中心数据集上联合训练单一模型,覆盖跨模态(MRI→CT、CBCT→CT)与同模态(MRI→MRI)任务。所有任务中,全体积处理优于基于块的方法,多任务模型性能等同于专用基线,却以一个模型替代了N个网络。关键突破:联合训练实现任务特定方法无法达到的能力——在未见解剖区域实现零样本泛化(SSIM达0.15),以及沿未直接监督路径进行组合式跨数据集翻译。结果表明,结合强体积先验与多任务训练是实现超越训练分布泛化的可扩展合成系统之路。代码已开源。

原文摘要 · Abstract (English)

Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitations: methods operate on 2D slices or 3D patches rather than whole volumes, and train a separate model for each translation task. Both stem from a single cause, the absence of a sufficiently strong volumetric prior, which forces generative models to learn anatomical appearance and cross-modality mapping simultaneously, an ill-posed problem at the scale of available paired datasets. We propose to decouple these objectives. A large-scale pretrained 3D variational autoencoder provides a compact latent representation of volumetric appearance, reducing translation to a conditional flow-matching problem. This compression makes whole-volume processing tractable, while a resolution-aware sampling strategy preserves native anatomical scale. We train a single model jointly across inter-modality (MRI$\to$CT, CBCT$\to$CT) and intra-modality (MRI$\to$MRI) tasks over three multi-center datasets. Across all tasks, whole-volume processing outperforms its patch-based counterpart, and the multi-task model matches task-specific baselines while replacing $N$ networks with one. Crucially, joint training unlocks capabilities inaccessible to task-specific approaches: zero-shot generalization to anatomical regions unseen during training, within 0.15 SSIM of the fully supervised model, and compositional cross-dataset translation along paths never directly supervised. These results suggest that combining a strong volumetric prior with multitask training is a scalable route toward synthesis systems that generalize beyond their training distribution. Code is available at https://github.com/arco-group/Whole-Volume-Latent-FM.

医学图像跨模态生成模型多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。