arXiv:2503.05319cs.CVcs.AI2025-03被引 4

通过解耦表示提升眼科疾病分级的多模态学习鲁棒性

Robust Multimodal Learning for Ophthalmic Disease Grading via Disentangled Representation

  • 引入自蒸馏机制,聚焦关键特征以减少冗余
  • 分离模态共性和特有表示,降低特征混淆
  • 适合医疗多模态数据稀缺场景下的模型优化

本文针对眼科医生依赖多模态数据提升诊断准确性的需求,指出真实场景中完整多模态数据因设备不足和隐私顾虑而罕见。传统深度学习方法虽在隐空间学习表征,但仍存在两大问题:(i) 复杂模态中任务无关的冗余信息(如大量切片)导致隐空间表征冗余;(ii) 多模态表征重叠,难以提取各模态独有特征。为此,作者提出本质点与解耦表示学习(EDRL)策略,将自蒸馏机制融入端到端框架,增强特征选择与解耦能力。具体而言,本质点表征学习模块筛选具有判别性的特征,提升疾病分级性能;解耦表征学习模块将多模态数据分解为模态共性和模态特有表示,降低特征纠缠,提升诊断鲁棒性与可解释性。在多个多模态眼科数据集上的实验表明,所提EDRL策略显著优于当前最先进方法。

原文摘要 · Abstract (English)

This paper discusses how ophthalmologists often rely on multimodal data to improve diagnostic accuracy. However, complete multimodal data is rare in real-world applications due to a lack of medical equipment and concerns about data privacy. Traditional deep learning methods typically address these issues by learning representations in latent space. However, the paper highlights two key limitations of these approaches: (i) Task-irrelevant redundant information (e.g., numerous slices) in complex modalities leads to significant redundancy in latent space representations. (ii) Overlapping multimodal representations make it difficult to extract unique features for each modality. To overcome these challenges, the authors propose the Essence-Point and Disentangle Representation Learning (EDRL) strategy, which integrates a self-distillation mechanism into an end-to-end framework to enhance feature selection and disentanglement for more robust multimodal learning. Specifically, the Essence-Point Representation Learning module selects discriminative features that improve disease grading performance. The Disentangled Representation Learning module separates multimodal data into modality-common and modality-unique representations, reducing feature entanglement and enhancing both robustness and interpretability in ophthalmic disease diagnosis. Experiments on multimodal ophthalmology datasets show that the proposed EDRL strategy significantly outperforms current state-of-the-art methods.

多模态学习眼科诊断解耦表征医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。