首个融合视频、评分与评语的手术技能评估基准,助力自动化手术培训评价。
SurgSkill-Bench: A Benchmark for Multimodal Surgical Skill Assessment

- 构建多模态手术视频-评分-评语数据集,支持视频与文本联合分析
- 在自定义二分法下,最优模型平均AUROC达0.88,文本评论显著提升评估效果
- 适合医学人工智能、手术训练系统研发者参考
手术技术能力的客观评估对医学生培训和结构化反馈至关重要,但当前流程仍依赖耗时的人工专家评审。现有自动化方法主要基于视觉输入,难以协同分析操作表现、结构化技能评分与评估者反馈。本文提出SurgSkill-Bench,一个初始的视频-评分-文本基准数据集,包含214个外科训练模拟视频、六维OSATS评分及专家自由文本评论。定义两种评估场景:仅用视频预测OSATS分数,以及后验式评论辅助预测(利用评论作为辅助信息)。采用代表性冻结视觉主干网络、内容自适应关键帧采样及简单视频-文本交叉注意力融合模块进行基线实验。内部视频级验证显示,内容自适应采样可提升视频仅预测性能,而评论在辅助设置中提供额外得分相关信号。在数据集特定中位数二分法下,最佳均值AUROC达到0.88。进一步讨论了数据规模、元数据完整性及评论辅助预测解释等评估限制。代码将在后续公开。
原文摘要 · Abstract (English)
Objective assessment of surgical technical skill is important for surgical training and structured feedback, but current workflows remain dependent on labor-intensive expert review. Existing automated approaches primarily focus on visual inputs and provide limited support for jointly studying operative performance, structured skill scores, and evaluator feedback. We introduce SurgSkill-Bench, an initial video-score-text benchmark-style dataset containing 214 surgical training simulation videos, six-dimensional OSATS scores, and expert free-text comments. We define two evaluation settings: video-only OSATS prediction for automated assessment and post hoc expert-comment-assisted prediction, where evaluator comments are available as auxiliary information. We provide controlled baseline experiments using representative frozen visual backbones, content-adaptive key-frame sampling, and a simple video-text co-attention fusion module. Under internal video-level validation, content-adaptive sampling improves video-only performance in this dataset, while evaluator comments provide additional score-related signal in the assisted setting. The best mean AUROC reaches 0.88 under dataset-specific median dichotomization. We further discuss evaluation constraints related to dataset scale, metadata completeness, and the interpretation of comment-assisted prediction. Code will be released publicly at a later date.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。