arXiv:2409.01459cs.CV2024-09被引 1

用3D大模型自动检测喉癌,准确率达92.4%。

3D-LSPTM: An Automatic Framework with 3D-Large-Scale Pretrained Model for Laryngeal Cancer Detection Using Laryngoscopic Videos

  • 基于C3D、TimeSformer、Video-Swin-Transformer的3D预训练模型微调
  • 在1109段喉镜视频上实现92.4%准确率、95.6%敏感度
  • 适合医疗AI开发者和耳鼻喉科医生参考

喉癌是耳鼻喉科高致死率恶性肿瘤,威胁人类健康。传统方法依赖医生手动观察喉镜视频,耗时且主观。本研究提出一种基于3D大模型的自动检测框架3D-LSPTM,利用中山大学附属第一医院伦理委员会批准的1,109段喉镜视频数据,采用C3D、TimeSformer、Video-Swin-Transformer等3D预训练模型,通过微调技术实现喉癌检测。大量实验表明,该框架表现优异,其中以Video-Swin-Transformer为骨干网络的3D-LSPTM达到92.4%准确率、95.6%敏感度、94.1%精确率和94.8%F₁值。

原文摘要 · Abstract (English)

Laryngeal cancer is a malignant disease with a high morality rate in otorhinolaryngology, posing an significant threat to human health. Traditionally larygologists manually visual-inspect laryngeal cancer in laryngoscopic videos, which is quite time-consuming and subjective. In this study, we propose a novel automatic framework via 3D-large-scale pretrained models termed 3D-LSPTM for laryngeal cancer detection. Firstly, we collect 1,109 laryngoscopic videos from the First Affiliated Hospital Sun Yat-sen University with the approval of the Ethics Committee. Then we utilize the 3D-large-scale pretrained models of C3D, TimeSformer, and Video-Swin-Transformer, with the merit of advanced featuring videos, for laryngeal cancer detection with fine-tuning techniques. Extensive experiments show that our proposed 3D-LSPTM can achieve promising performance on the task of laryngeal cancer detection. Particularly, 3D-LSPTM with the backbone of Video-Swin-Transformer can achieve 92.4% accuracy, 95.6% sensitivity, 94.1% precision, and 94.8% F_1.

喉癌检测3D模型医学影像视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。