构建首个统一的手术视频分析基准,推动自动化手术决策与技能评估
SurgBench: A Unified Large-Scale Benchmark for Surgical Video Analysis
- 构建涵盖22种术式、11个专科的5300万帧预训练数据集
- 在72项细粒度任务上验证现有模型泛化能力差,预训练后性能显著提升
- 适合医疗AI研究者、手术机器人开发者及医学影像团队使用
手术视频理解对实现术中自动化决策、技能评估和术后质量改进至关重要。然而,由于缺乏大规模、多样化的预训练数据集和系统性评估标准,手术视频基础模型的发展受到制约。本文提出SurgBench,一个统一的手术视频基准框架,包含预训练数据集SurgBench-P和评估基准SurgBench-E。SurgBench-P覆盖22种手术流程、11个专科,包含5300万帧视频;SurgBench-E在6类任务(阶段分类、摄像机运动、器械识别、疾病诊断、动作分类、器官检测)下涵盖72项细粒度评估任务。大量实验表明,现有视频基础模型在不同任务间泛化能力有限,而基于SurgBench-P预训练可显著提升性能并增强跨领域泛化能力,适用于未见术式与模态。数据与代码可申请获取。
原文摘要 · Abstract (English)
Surgical video understanding is pivotal for enabling automated intraoperative decision-making, skill assessment, and postoperative quality improvement. However, progress in developing surgical video foundation models (FMs) remains hindered by the scarcity of large-scale, diverse datasets for pretraining and systematic evaluation. In this paper, we introduce \textbf{SurgBench}, a unified surgical video benchmarking framework comprising a pretraining dataset, \textbf{SurgBench-P}, and an evaluation benchmark, \textbf{SurgBench-E}. SurgBench offers extensive coverage of diverse surgical scenarios, with SurgBench-P encompassing 53 million frames across 22 surgical procedures and 11 specialties, and SurgBench-E providing robust evaluation across six categories (phase classification, camera motion, tool recognition, disease diagnosis, action classification, and organ detection) spanning 72 fine-grained tasks. Extensive experiments reveal that existing video FMs struggle to generalize across varied surgical video analysis tasks, whereas pretraining on SurgBench-P yields substantial performance improvements and superior cross-domain generalization to unseen procedures and modalities. Our dataset and code are available upon request.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。