通过时空注意力与边缘引导特征学习,提升结肠段识别准确率
ST-ColoNet: Spatio-Temporal Colon Segment Recognition via Hybrid Attention and Edge-Guided Feature Learning
- 融合时空注意力与边缘引导机制,捕捉视频中空间和时间特征
- 在自建数据集上达81.0%准确率和70.7%F1分数,显著优于现有方法
- 适合医学影像分析、内窥镜视频理解方向的研究者参考
结肠镜视频中的结肠段识别是众多下游任务的关键,但现有自动方法仅使用图像而未充分利用时序信息,导致性能不佳。此外,相关公开视频数据集稀缺。为此,我们构建并发布了专用于结肠段识别的标注数据集。同时提出两阶段深度学习框架ST-ColoNet,包含利用度量学习优化边缘引导空间特征提取的Colorlaus模块,以及结合三种自注意力模式以更好近似长序列全自注意力的Full-T模块,从而优化时序特征聚合。大量消融实验表明,该框架在结肠段识别任务上达到81.0%准确率和70.7%F1分数,远超当前最优方法。
原文摘要 · Abstract (English)
Colo-segment recognition in colonoscopy videos is a key requirement for many downstream tasks, but existing automatic recognition methods only use colonoscopy images without fully exploiting the use of temporal information, leading to poor performance. Additionally, relevant public video-based datasets are in scarcity. To tackle this problem, we curate and release a labeled dataset specifically for the task of colo-segment recognition. In addition, we propose a two-stage deep learning-based framework, Colo-Segment Recognition via SpatioTemporal Network (ST-ColoNet), for the task of colo-segment recognition from colonoscopy videos which includes the Colorlaus module that uses metric learning to optimize edge-mediated spatial feature extraction, as well as the Full-Temp module which combines three self-attention patterns to better approximate full self-attention on long colonoscopy sequences and optimize temporal feature aggregation. Through extensive ablation experiments, we show that our framework is capable of achieving state-of-the-art performance on the task of colo-segment recognition, achieving an accuracy of 81.0% and F1-score of 70.7%, which is a tremendous improvement over state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。