用62张3D MRI训练出能对齐医学影像与表格数据的专用模型。
Revisiting CLIP: Efficient Alignment of 3D MRI and Tabular Data using Domain-Specific Foundation Models
- 用3D基础模型替代传统CLIP,通过嵌入累积策略稳定训练
- 仅用62例MRI即实现3D影像与表格数据的有效对齐
- 适合医疗多模态研究者,尤其关注小样本场景
多模态模型需共享对齐的嵌入空间。但主流的CLIP方法需要大量样本,且不原生支持3D或表格数据,而这在医疗领域至关重要。为此,我们重新审视CLIP式对齐,训练一个专用3D基础模型作为图像编码器,并证明仅需62例MRI扫描即可实现模态对齐。该方法依赖于一种简单的嵌入累积策略,可扩展批次间的负样本数量以稳定训练。我们系统评估了骨干网络和损失函数等设计选择,并在零样本分类与图像检索任务上测试了所提方法。尽管零样本图像检索仍具挑战,零样本分类结果表明该方法能有效对齐3D MRI与表格数据的表征。
原文摘要 · Abstract (English)
Multi-modal models require aligned, shared embedding spaces. However, common CLIP-based approaches need large amounts of samples and do not natively support 3D or tabular data, both of which are crucial in the medical domain. To address these issues, we revisit CLIP-style alignment by training a domain-specific 3D foundation model as an image encoder and demonstrate that modality alignment is feasible with only 62 MRI scans. Our approach is enabled by a simple embedding accumulation strategy required for training in 3D, which scales the amount of negative pairs across batches in order to stabilize training. We perform a thorough evaluation of various design choices, including the choice of backbone and loss functions, and evaluate the proposed methodology on zero-shot classification and image-retrieval tasks. While zero-shot image-retrieval remains challenging, zero-shot classification results demonstrate that the proposed approach can meaningfully align the representations of 3D MRI with tabular data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。