构建2.4万段内镜手术视频标注数据集,助力自动索引与分析。
TEMSET-24K: Densely Annotated Dataset for Indexing Multipart Endoscopic Videos using Surgical Timeline Segmentation
- 设计分层标签体系,精细标注手术阶段、任务与动作
- 模型在关键阶段识别准确率高达0.99,F1达0.99
- 开源数据集适合做手术流程自动化研究的团队使用
内镜手术视频的索引在手术数据科学中至关重要,是开展回顾性分析和临床能力评估的基础。然而,当前视频分析仍依赖人工标注,效率低下。尽管计算机视觉尤其是深度学习具备自动化潜力,但受限于缺乏公开可用的密集标注手术数据集。为此,我们提出TEMSET-24K,一个包含24,306段经肛门内镜微创手术(TEMS)视频微片段的开源数据集。每段视频均由临床专家采用新型分层标注体系(阶段、任务、动作三元组)进行细致标注,完整捕捉复杂的手术流程。为验证该数据集,我们对多种基于Transformer的深度学习模型进行了基准测试。仿真评估显示,关键阶段如准备与缝合的识别准确率最高达0.99,F1分数最高达0.99。采用ConvNeXt、ViT和SWIN V2编码器的STALNet模型,在代表性阶段上表现稳定。TEMSET-24K为手术数据科学提供了关键基准,推动了该领域前沿解决方案的发展。
原文摘要 · Abstract (English)
Indexing endoscopic surgical videos is vital in surgical data science, forming the basis for systematic retrospective analysis and clinical performance evaluation. Despite its significance, current video analytics rely on manual indexing, a time-consuming process. Advances in computer vision, particularly deep learning, offer automation potential, yet progress is limited by the lack of publicly available, densely annotated surgical datasets. To address this, we present TEMSET-24K, an open-source dataset comprising 24,306 trans-anal endoscopic microsurgery (TEMS) video micro-clips. Each clip is meticulously annotated by clinical experts using a novel hierarchical labeling taxonomy encompassing phase, task, and action triplets, capturing intricate surgical workflows. To validate this dataset, we benchmarked deep learning models, including transformer-based architectures. Our in silico evaluation demonstrates high accuracy (up to 0.99) and F1 scores (up to 0.99) for key phases like Setup and Suturing. The STALNet model, tested with ConvNeXt, ViT, and SWIN V2 encoders, consistently segmented well-represented phases. TEMSET-24K provides a critical benchmark, propelling state-of-the-art solutions in surgical data science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。