为内窥镜视频设计过滤方案,提升自监督学习效率与诊断准确率。
A filtering scheme for confocal laser endomicroscopy (CLE)-video sequences for self-supervised learning
- 提出视频过滤方法,降低内窥镜序列中的帧间冗余。
- 在两个肿瘤数据集上,模型准确率达67.48%和73.52%,显著优于基线。
- 过滤后训练时间减少67%,适合资源受限的医学影像自监督场景。
共聚焦激光内窥镜(CLE)是一种非侵入性实时成像技术,可用于黏膜微结构的在体、在位分析。然而,其图像对非专业医生而言难以解读,引入机器学习可辅助诊断,但受限于与组织病理学关联的CLE序列稀缺,易导致模型过拟合。为此,可采用自监督学习(SSL)在大规模无标注数据上预训练。由于CLE为视频模态,帧间相关性高,导致SSL训练数据分布不均衡。本文提出一种针对CLE视频序列的过滤机制,降低数据冗余,提升SSL训练收敛速度与效率。采用四种主流网络及基于小规模视觉变换器的教师-学生网络进行评估。在鼻窦肿瘤与皮肤鳞状细胞癌数据集上的下游任务中,过滤后的SSL预训练模型分别达到67.48%和73.52%的最高测试准确率,显著优于非SSL基线。结果表明,SSL适用于CLE预训练;且所提过滤方法可有效提升自监督训练效率,训练时间减少67%。
原文摘要 · Abstract (English)
Confocal laser endomicroscopy (CLE) is a non-invasive, real-time imaging modality that can be used for in-situ, in-vivo imaging and the microstructural analysis of mucous structures. The diagnosis using CLE is, however, complicated by images being hard to interpret for non-experienced physicians. Utilizing machine learning as an augmentative tool would hence be beneficial, but is complicated by the shortage of histopathology-correlated CLE imaging sequences with respect to the plurality of patterns in this domain, leading to overfitting of machine learning models. To overcome this, self-supervised learning (SSL) can be employed on larger unlabeled datasets. CLE is a video-based modality with high inter-frame correlation, leading to a non-stratified data distribution for SSL training. In this work, we propose a filter functionality on CLE video sequences to reduce the dataset redundancy in SSL training and improve SSL training convergence and training efficiency. We use four state-of-the-art baseline networks and a SSL teacher-student network with a vision transformer small backbone for the evaluation. These networks were evaluated on downstream tasks for a sinonasal tumor dataset and a squamous cell carcinoma of the skin dataset. On both datasets, we found the highest test accuracy on the filtered SSL-pretrained model, with 67.48% and 73.52%, both considerably outperforming their non-SSL baselines. Our results show that SSL is an effective method for CLE pretraining. Further, we show that our proposed CLE video filter can be utilized to improve training efficiency in self-supervised scenarios, resulting in a reduction of 67% in training time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。