统一医学影像表征学习,支持多任务灵活组合。
Universal Medical Image Representation Learning with Compositional Decoders
- 拆解-重构双解码器设计,实现像素与语义输出协同预测。
- 在8个数据集上达顶尖性能,零样本与100样本迁移能力强。
- 适合医学影像多任务研究者,尤其关注跨任务泛化场景。
视觉语言模型推动了通用模型的发展,但在医学影像领域受限于特定功能需求和数据稀缺。现有通用模型通常采用任务专用分支和头部,限制了共享特征空间与模型灵活性。为此,我们提出一种分解-组合式通用医学影像范式(UniMed),支持各级任务。首先设计分解解码器,基于输入队列预测像素与语义两类输出;进而引入组合解码器,统一输入输出空间,并将不同层级的任务标注转化为离散标记格式。两者耦合设计使模型能灵活组合任务并实现相互增益。此外,联合表征学习策略巧妙利用大量无标签数据与无监督损失,实现高效单阶段预训练,提升鲁棒性。实验表明,UniMed在8个数据集上三种任务均达到最先进水平,具备强大的零样本与100样本迁移能力。代码与训练模型将在论文录用后公开。
原文摘要 · Abstract (English)
Visual-language models have advanced the development of universal models, yet their application in medical imaging remains constrained by specific functional requirements and the limited data. Current general-purpose models are typically designed with task-specific branches and heads, which restricts the shared feature space and the flexibility of model. To address these challenges, we have developed a decomposed-composed universal medical imaging paradigm (UniMed) that supports tasks at all levels. To this end, we first propose a decomposed decoder that can predict two types of outputs -- pixel and semantic, based on a defined input queue. Additionally, we introduce a composed decoder that unifies the input and output spaces and standardizes task annotations across different levels into a discrete token format. The coupled design of these two components enables the model to flexibly combine tasks and mutual benefits. Moreover, our joint representation learning strategy skilfully leverages large amounts of unlabeled data and unsupervised loss, achieving efficient one-stage pretraining for more robust performance. Experimental results show that UniMed achieves state-of-the-art performance on eight datasets across all three tasks and exhibits strong zero-shot and 100-shot transferability. We will release the code and trained models upon the paper's acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。