构建首个大规模腹腔镜手术技能评估数据集,支持自动评分与错误识别。
A benchmark for video-based laparoscopic skill analysis and assessment
- 构建1270段双目视频数据集,涵盖4类基础训练任务。
- 每段视频含多方评分及任务特异性错误标签,标注一致性强。
- 提供预设划分数据集和基准模型,便于新方法对比评估。
腹腔镜手术是一项复杂技术,需长期训练。近年来,深度学习在自动视频化手术技能评估方面展现出潜力,但受限于现有标注数据集规模较小。为此,我们提出腹腔镜技能分析与评估(LASANA)数据集,包含1270段双目视频,覆盖4种基础训练任务。每段视频均经三位独立评审员标注结构化技能评分,并附有任务相关的二值错误标签。多数视频来自真实培训课程,体现学员技能的自然差异。为促进现有及新型视频化技能评估与错误识别方法的基准测试,我们为每项任务提供预定义数据划分。此外,还给出了基于深度学习模型的基线结果,作为未来研究的参考。
原文摘要 · Abstract (English)
Laparoscopic surgery is a complex surgical technique that requires extensive training. Recent advances in deep learning have shown promise in supporting this training by enabling automatic video-based assessment of surgical skills. However, the development and evaluation of deep learning models is currently hindered by the limited size of available annotated datasets. To address this gap, we introduce the Laparoscopic Skill Analysis and Assessment (LASANA) dataset, comprising 1270 stereo video recordings of four basic laparoscopic training tasks. Each recording is annotated with a structured skill rating, aggregated from three independent raters, as well as binary labels indicating the presence or absence of task-specific errors. The majority of recordings originate from a laparoscopic training course, thereby reflecting a natural variation in the skill of participants. To facilitate benchmarking of both existing and novel approaches for video-based skill assessment and error recognition, we provide predefined data splits for each task. Furthermore, we present baseline results from a deep learning model as a reference point for future comparisons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。