融合影像与临床数据的可解释模型提升口腔癌前病变检测准确率
Explainable Multimodal Deep Learning Integrating Imaging and Clinical Data for Oral Potentially Malignant Disorder Detection

- 整合白光与自荧光影像及标准化临床问卷,实现多模态联合分析
- 在真实筛查场景下达到AUC 0.952,对视觉不明显的病灶识别更优
- 通过可解释性分析揭示临床因素对诊断的关键贡献,适合临床部署
口腔潜在恶性病变(OPMD)是口腔癌的重要前期表现,但因其表型异质性强且与良性病变重叠,临床检测仍具挑战。尽管基于图像的深度学习在自动化筛查中展现出潜力,但在实际诊疗中,仅依赖视觉信息仍不足,还需结合患者个体风险因素。为此,我们构建了M2-OPMDNet,一种融合共配准的白光与自荧光口内影像及结构化临床信息的多模态深度学习框架。通过定制化问卷标准化采集临床相关风险因素与症状,并与图像特征集成。采用前瞻性收集的真实世界筛查数据集,评估了多种图像编码器(包括传统CNN与基于基础模型的架构)。利用SHapley Additive exPlanations(SHAP)分析模型可解释性,量化各特征与模态的贡献。M2-OPMDNet取得AUC 0.952,优于单一模态方法,尤其提升了对视觉不明显病灶的识别能力。SHAP分析显示,结构化临床变量在风险评估中贡献显著,补充了影像特征。结果表明,将白光与自荧光成像结合结构化临床数据的可解释多模态学习,可提供准确、透明且符合临床逻辑的OPMD检测。M2-OPMDNet为真实世界口腔癌筛查与决策支持提供可扩展的解决方案。
原文摘要 · Abstract (English)
Oral potentially malignant disorders (OPMDs) are critical precursors to oral cancer, yet clinical detection remains challenging because of substantial phenotypic heterogeneity and overlap with benign conditions. Although image-based deep learning shows promise for automated screening, visual information alone may be insufficient in real-world settings, where diagnostic decisions also rely on patient-specific risk factors. We developed M2-OPMDNet, a multimodal deep learning framework that integrates co-registered white-light and autofluorescence intraoral images with structured clinical information for OPMD detection. A customized questionnaire was designed to capture clinically relevant risk factors and symptoms in a standardized, reproducible format for integration with image-derived features. Multiple image encoders, including conventional convolutional neural networks and foundation model-based architectures, were evaluated using a prospectively collected dataset reflecting real-world screening conditions. Model interpretability was assessed using SHapley Additive exPlanations (SHAP) to quantify feature- and modality-level contributions. M2-OPMDNet achieved an AUC of 0.952, outperforming unimodal approaches and showing improved performance for visually subtle lesions. SHAP analysis demonstrated that structured clinical variables contributed substantially to risk estimation and complemented imaging features. These results demonstrate that explainable multimodal learning combining white-light and autofluorescence imaging with structured clinical data can provide accurate, transparent, and clinically grounded OPMD detection. M2-OPMDNet offers a scalable framework for real-world oral cancer screening and decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。