随机融合机制提升多视角医学影像诊断准确率
Random Token Fusion for Multi-View Medical Diagnosis
- 训练时引入随机令牌融合,增强特征融合的鲁棒性
- 在乳腺钼靶和胸片数据集上显著提升诊断性能
- 无需额外计算开销,适合构建下一代医学大模型
在多视角医学诊断中,基于深度学习的模型常通过融合不同成像视角的信息来提升诊断效果。然而,现有方法易过拟合且过度依赖视图特异性特征,可能导致平凡解。本文提出随机令牌融合(RTF)技术,利用视觉变换器进行多视角医学图像分析。通过在训练阶段引入随机性,有效缓解过拟合问题,提升模型鲁棒性与准确性,且推理阶段无额外开销。我们在标准乳腺钼靶和胸部X光数据集上验证了该方法,实验表明RTF能持续改进现有融合方法的性能,为新一代多视角医学基础模型的发展提供新路径。
原文摘要 · Abstract (English)
In multi-view medical diagnosis, deep learning-based models often fuse information from different imaging perspectives to improve diagnostic performance. However, existing approaches are prone to overfitting and rely heavily on view-specific features, which can lead to trivial solutions. In this work, we introduce Random Token Fusion (RTF), a novel technique designed to enhance multi-view medical image analysis using vision transformers. By integrating randomness into the feature fusion process during training, RTF addresses the issue of overfitting and enhances the robustness and accuracy of diagnostic models without incurring any additional cost at inference. We validate our approach on standard mammography and chest X-ray benchmark datasets. Through extensive experiments, we demonstrate that RTF consistently improves the performance of existing fusion methods, paving the way for a new generation of multi-view medical foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。