arXiv:2602.17901eess.IVcs.CV2026-02

通过解耦医学影像内容与风格,实现跨中心3D影像合成与分析统一建模。

MeDUET: Disentangled Unified Pretraining for 3D Medical Image Synthesis and Analysis

  • 在变分自编码器潜空间中解耦器官结构与扫描风格,提升模型泛化能力。
  • 在多中心数据上实现更高保真度的影像生成和更快的训练收敛速度。
  • 适合需要跨设备、跨模态医学影像建模的研究者与临床应用开发者。

自监督学习(SSL)与扩散模型分别推动了高维3D视觉数据的表征学习与生成建模,但常作为独立范式发展。在多源异构条件下,其统一面临挑战:解剖结构需保持一致以支持分析,而不同中心的采集风格差异影响生成质量。本文提出MeDUET,一种基于变分自编码器潜空间的3D医学图像解耦统一预训练框架。将统一预训练建模为经验因子可识别问题,旨在学习域不变的解剖内容因子与域特定的外观风格因子。为增强因子分离,先使用令牌解混搭配标准对抗域正则化建立基础的内容-风格专属性,再引入混合因子令牌蒸馏与交换不变四元组对比损失,减少混合区域因子泄漏,并通过因子级不变性与可区分性组织因子空间。利用所学因子,MeDUET在合成与分析任务间有效迁移,生成更高质量图像,收敛更快,可控性更强;同时在多种数据集、任务与模态下达到竞争性或更优的域泛化与标签效率表现。结果表明,多源异构性可作为有效监督信号,解耦提供统一3D医学影像合成与分析的有效接口。代码已开源。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) and diffusion models have respectively advanced representation learning and generative modeling for high-dimensional 3D visual data, yet they are often developed as separate paradigms. Their unification remains challenging under multi-source heterogeneity, as anatomical content must be preserved for analysis while acquisition-related style varies across centers and affects synthesis. In this paper, we propose MeDUET, a 3D Medical image Disentangled UnifiEd PreTraining framework in the variational autoencoder latent space. MeDUET formulates unified pretraining as an empirical factor identifiability problem, aiming to learn domain-invariant content factors for anatomy and domain-specific style factors for appearance. To improve factor separation, MeDUET first uses token demixing with a standard adversarial domain regularizer to establish basic content-style specialization, and further introduces Mixed Factor Token Distillation and Swap-invariance Quadruplet Contrast to reduce mixed-region factor leakage and organize factor spaces with factor-wise invariance and discriminability. With these learned factors, MeDUET transfers effectively to both synthesis and analysis, yielding higher fidelity, faster convergence, and better controllability for synthesis, while achieving competitive or superior domain generalization and label efficiency on diverse datasets, tasks, and modalities. Overall, MeDUET shows that multi-source heterogeneity can serve as useful supervision, with disentanglement providing an effective interface for unifying 3D medical image synthesis and analysis. Our code is available at https://github.com/JK-Liu7/MeDUET.

3D医学影像解耦表征统一建模自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。