构建多视角手术手势与错误数据集,助力自动化医训评估
SHANDS: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical Training
- 五视角拍摄52人完成标准手术操作,标注15种手势和8类错误
- 提出单/多视角及跨视角通用性评测协议,验证主流模型性能
- 数据集公开,适合医疗视觉与智能训练系统研发者使用
在医学教学中,外科技能培养依赖专家评估,但成本高、难扩展且资源集中。现有自动化评估受限于缺乏包含真实学员错误和多视角变化的数据集。为此,我们构建了大型多视角视频数据集Surgical-Hands(SHands),通过五个RGB摄像头从互补视角记录52名参与者(20位专家,32名学员)完成线性切开与缝合,每人执行三轮标准化操作。视频逐帧标注15种手势基元,并包含经验证的8类学员错误分类体系,支持手势识别与错误检测。我们定义了单视角、多视角及跨视角泛化标准评测协议,并对主流深度学习模型进行基准测试。SHands已公开发布,旨在推动基于临床知识的鲁棒、可扩展智能外科训练系统发展。
原文摘要 · Abstract (English)
In surgical training for medical students, proficiency development relies on expert-led skill assessment, which is costly, time-limited, difficult to scale, and its expertise remains confined to institutions with available specialists. Automated AI-based assessment offers a viable alternative, but progress is constrained by the lack of datasets containing realistic trainee errors and the multi-view variability needed to train robust computer vision approaches. To address this gap, we present Surgical-Hands (SHands), a large-scale multi-view video dataset for surgical hand-gesture and error recognition for medical training. \textsc{SHands} captures linear incision and suturing using five RGB cameras from complementary viewpoints, performed by 52 participants (20 experts and 32 trainees), each completing three standardized trials per procedure. The videos are annotated at the frame level with 15 gesture primitives and include a validated taxonomy of 8 trainee error types, enabling both gesture recognition and error detection. We further define standardized evaluation protocols for single-view, multi-view, and cross-view generalization, and benchmark state-of-the-art deep learning models on the dataset. SHands is publicly released to support the development of robust and scalable AI systems for surgical training grounded in clinically curated domain knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。