用少样本学习优化肠镜图像分类,提升早期肠癌检测准确率
Lightweight Relational Embedding in Task-Interpolated Few-Shot Networks for Enhanced Gastrointestinal Disease Classification
- 基于任务插值与关系嵌入的轻量模型,增强对相似肠镜帧的区分能力
- 在Kvasir数据集上达90.1%准确率,F1-score 0.891,优于现有方法
- 适合医疗图像分析、少样本学习及内窥镜辅助诊断场景
传统诊断方法如结肠镜检查虽关键但具侵入性,且依赖高质量内镜图像。早期发现结直肠癌(CRC)对提高生存率至关重要,但图像质量不足或重复模式导致判别困难。为此,我们提出一种基于少样本学习的新深度网络,包含定制特征提取器、任务插值、关系嵌入及双层路由注意力机制。该模型可快速适应未见细粒度图像模式,任务插值从不同视角人工扩充图像。关系嵌入捕捉图像内部关键特征与连续帧间动态变化,克服了卷积神经网络(CNNs)局限。轻量注意力机制聚焦相关区域,提升分析效率。在多数据集训练下显著增强泛化性与鲁棒性。在Kvasir数据集上,模型实现90.1%准确率、0.845精确率、0.942召回率和0.891 F1分数,超越当前最优方法,为优化侵入性结肠镜检查下的CRC检测提供高效图像分析方案。
原文摘要 · Abstract (English)
Traditional diagnostic methods like colonoscopy are invasive yet critical tools necessary for accurately diagnosing colorectal cancer (CRC). Detection of CRC at early stages is crucial for increasing patient survival rates. However, colonoscopy is dependent on obtaining adequate and high-quality endoscopic images. Prolonged invasive procedures are inherently risky for patients, while suboptimal or insufficient images hamper diagnostic accuracy. These images, typically derived from video frames, often exhibit similar patterns, posing challenges in discrimination. To overcome these challenges, we propose a novel Deep Learning network built on a Few-Shot Learning architecture, which includes a tailored feature extractor, task interpolation, relational embedding, and a bi-level routing attention mechanism. The Few-Shot Learning paradigm enables our model to rapidly adapt to unseen fine-grained endoscopic image patterns, and the task interpolation augments the insufficient images artificially from varied instrument viewpoints. Our relational embedding approach discerns critical intra-image features and captures inter-image transitions between consecutive endoscopic frames, overcoming the limitations of Convolutional Neural Networks (CNNs). The integration of a light-weight attention mechanism ensures a concentrated analysis of pertinent image regions. By training on diverse datasets, the model's generalizability and robustness are notably improved for handling endoscopic images. Evaluated on Kvasir dataset, our model demonstrated superior performance, achieving an accuracy of 90.1\%, precision of 0.845, recall of 0.942, and an F1 score of 0.891. This surpasses current state-of-the-art methods, presenting a promising solution to the challenges of invasive colonoscopy by optimizing CRC detection through advanced image analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。