arXiv:2511.21582cs.CV2025-11

通过多模态融合与数据增强提升口腔癌病变识别准确率

Data-Augmented Multimodal Feature Fusion for Multiclass Visual Recognition of Oral Cancer Lesions

  • 结合临床信息与图像数据,用数据增强提升特征多样性
  • 四类识别准确率达54.97%,优于传统单模态模型
  • 适合开发沉浸式虚拟现实辅助诊断系统

口腔癌常因与其它病灶相似而被延误诊断。现有基于深度学习的辅助诊断研究受限于小样本、不平衡数据集及单一模态特征,影响模型在真实临床环境中的泛化能力。为此,本研究提出一种数据增强驱动的多模态特征融合框架,集成于视觉识别(VR)辅助口腔癌识别系统中。方法结合大量以数据为中心的增强与临床-图像双模态表征,提升模型鲁棒性并降低诊断歧义。采用分层训练流程与EfficientNetV2 B1主干网络,增强特征多样性,缓解数据不平衡,强化多模态嵌入表达。实验表明,该框架在二分类任务中达到82.57%准确率,三分类为65.13%,四分类为54.97%,显著优于传统单流CNN模型。结果验证了多模态融合与策略性增强在早期口腔癌病灶可靠识别中的有效性,并为沉浸式虚拟现实临床决策支持工具奠定基础。

原文摘要 · Abstract (English)

Oral cancer is frequently diagnosed at later stages due to its similarity to other lesions. Existing research on computer aided diagnosis has made progress using deep learning; however, most approaches remain limited by small, imbalanced datasets and a dependence on single-modality features, which restricts model generalization in real-world clinical settings. To address these limitations, this study proposes a novel data-augmentation driven multimodal feature-fusion framework integrated within a (Vision Recognition)VR assisted oral cancer recognition system. Our method combines extensive data centric augmentation with fused clinical and image-based representations to enhance model robustness and reduce diagnostic ambiguity. Using a stratified training pipeline and an EfficientNetV2 B1 backbone, the system improves feature diversity, mitigates imbalance, and strengthens the learned multimodal embeddings. Experimental evaluation demonstrates that the proposed framework achieves an overall accuracy of 82.57 percent on 2 classes, 65.13 percent on 3 classes, and 54.97 percent on 4 classes, outperforming traditional single stream CNN models. These results highlight the effectiveness of multimodal feature fusion combined with strategic augmentation for reliable early oral cancer lesion recognition and serve as a foundation for immersive VR based clinical decision support tools.

口腔癌多模态融合数据增强视觉识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。