用优化网络实时识别白内障手术视频中的器械,准确率超73%。
Identifying Surgical Instruments in Pedagogical Cataract Surgery Videos through an Optimized Aggregation Network
- 基于YOLOv9改进,引入PGI和新聚合结构解决信息瓶颈。
- 在615张图、10类器械上达73.74% mAP(IoU=0.5)。
- 适合眼科教学视频分析与手术辅助系统研发者。
教学用白内障手术视频对眼科医生和学员反复观察手术细节至关重要。本文提出一种深度学习模型,用于实时识别此类视频中的手术器械,使用从公开来源抓取的自定义数据集。受YOLOv9架构启发,模型采用可编程梯度信息(PGI)机制和新型通用优化高效层聚合网络(Go-ELAN),以缓解信息瓶颈问题,提升高非极大值抑制交并比(NMS IoU)下的最小平均精度(mAP)。Go-ELAN YOLOv9模型在包含615张图像、10类器械的数据集上,于IoU=0.5时达到73.74%的mAP,优于YOLO v5、v7、v8、v9原版、Laptool和DETR,验证了该模型的有效性。
原文摘要 · Abstract (English)
Instructional cataract surgery videos are crucial for ophthalmologists and trainees to observe surgical details repeatedly. This paper presents a deep learning model for real-time identification of surgical instruments in these videos, using a custom dataset scraped from open-access sources. Inspired by the architecture of YOLOV9, the model employs a Programmable Gradient Information (PGI) mechanism and a novel Generally-Optimized Efficient Layer Aggregation Network (Go-ELAN) to address the information bottleneck problem, enhancing Minimum Average Precision (mAP) at higher Non-Maximum Suppression Intersection over Union (NMS IoU) scores. The Go-ELAN YOLOV9 model, evaluated against YOLO v5, v7, v8, v9 vanilla, Laptool and DETR, achieves a superior mAP of 73.74 at IoU 0.5 on a dataset of 615 images with 10 instrument classes, demonstrating the effectiveness of the proposed model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。