用多模态训练知识蒸馏提升视觉模型诊断能力
On the effectiveness of multimodal privileged knowledge distillation in two vision transformer based diagnostic applications
- 训练时用文本/表格数据辅助视觉模型,推理时仅需图像
- 在胸片和乳腺影像上显著提升注意力定位病灶能力
- 适合临床部署中模态缺失场景的模型优化
深度学习模型在临床应用中常需融合图像、文本和结构化数据等多模态信息以做出可靠决策。然而,推理时并非所有模态都可用。本文提出多模态特权知识蒸馏(MMPKD),利用仅在训练阶段可用的额外模态来指导单模态视觉模型。具体地,针对胸部X光(MIMIC-CXR)采用基于文本的教师模型,针对乳腺影像(CBIS-DDSM)采用基于表格元数据的教师模型,将知识蒸馏至视觉变压器学生模型。实验表明,MMPKD可显著提升注意力图在零样本条件下定位输入图像中感兴趣区域的能力,但该效果不具备跨领域泛化性,与以往研究结论相反。
原文摘要 · Abstract (English)
Deploying deep learning models in clinical practice often requires leveraging multiple data modalities, such as images, text, and structured data, to achieve robust and trustworthy decisions. However, not all modalities are always available at inference time. In this work, we propose multimodal privileged knowledge distillation (MMPKD), a training strategy that utilizes additional modalities available solely during training to guide a unimodal vision model. Specifically, we used a text-based teacher model for chest radiographs (MIMIC-CXR) and a tabular metadata-based teacher model for mammography (CBIS-DDSM) to distill knowledge into a vision transformer student model. We show that MMPKD can improve the resulting attention maps' zero-shot capabilities of localizing ROI in input images, while this effect does not generalize across domains, as contrarily suggested by prior research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。