arXiv:2412.18124cs.CV2024-12被引 2

用图文融合模型提升声带癌早期诊断准确率

VisionLLM-based Multimodal Fusion Network for Glottic Carcinoma Early Detection

  • 结合图像与文本信息,通过视觉大模型提取多模态特征
  • 在5799对图文数据上实现当前最优检测性能
  • 适合医学影像分析与多模态算法研究者参考

声带癌的早期检测对改善患者预后至关重要,可实现及时干预、保留声功能,并显著降低肿瘤进展和转移风险。然而,声带癌与声带异型增生在形态上相似,导致检测准确率不理想。为此,本文提出一种基于视觉大语言模型(VisionLLM)的多模态融合网络MMGC-Net,用于声带癌检测。通过整合图像与文本模态,多模态模型能够捕捉互补信息,提升预测准确性与鲁棒性。本文收集了来自中山大学附属第一医院的真实声带癌数据集SYSU1H,包含5,799对图像-文本数据。采用图像编码器与额外Q-Former提取视觉嵌入,利用大型语言模型Meta AI(Llama3)获取文本嵌入,并通过喉部特征融合模块实现图像与文本特征的综合集成,从而提升声带癌识别能力。在SYSU1H数据集上的大量实验表明,MMGC-Net达到当前最优性能,优于以往多模态模型。

原文摘要 · Abstract (English)

The early detection of glottic carcinoma is critical for improving patient outcomes, as it enables timely intervention, preserves vocal function, and significantly reduces the risk of tumor progression and metastasis. However, the similarity in morphology between glottic carcinoma and vocal cord dysplasia results in suboptimal detection accuracy. To address this issue, we propose a vision large language model-based (VisionLLM-based) multimodal fusion network for glottic carcinoma detection, known as MMGC-Net. By integrating image and text modalities, multimodal models can capture complementary information, leading to more accurate and robust predictions. In this paper, we collect a private real glottic carcinoma dataset named SYSU1H from the First Affiliated Hospital of Sun Yat-sen University, with 5,799 image-text pairs. We leverage an image encoder and additional Q-Former to extract vision embeddings and the Large Language Model Meta AI (Llama3) to obtain text embeddings. These modalities are then integrated through a laryngeal feature fusion block, enabling a comprehensive integration of image and text features, thereby improving the glottic carcinoma identification performance. Extensive experiments on the SYSU1H dataset demonstrate that MMGC-Net can achieve state-of-the-art performance, which is superior to previous multimodal models.

医学影像多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。