构建首个大规模手术动作数据集,实现跨专业手术动作识别。
Generalized Recognition of Basic Surgical Actions Enables Skill Assessment and Vision-Language-Model-based Surgical Planning
- 基于11000+视频片段构建6类外科手术的10种基础动作数据集
- 模型在前列腺切除等任务中实现精准技能评估与可解释性规划
- 适合研究手术智能、医学AI融合与临床辅助系统开发者
人工智能、医学影像与大语言模型有望重塑外科实践、培训与自动化。理解并建模基本手术动作(BSA)——任何手术的基本单元——是推动该领域发展的关键。本文提出一个包含6个外科专科、10种基本动作、超过11,000个视频片段的BSA数据集,为迄今最大规模。基于此,我们开发了一个通用基础模型,可实现基础动作的广泛识别。实验表明该模型在不同术式和解剖部位的数据集上均表现出稳健的跨专科性能。进一步,我们展示了该基础模型的下游应用:在前列腺切除术中结合领域知识进行技能评估,在胆囊切除术和肾切除术中利用大视觉-语言模型实现动作规划。多位国际外科医生对语言模型生成的可解释性规划文本进行了评估,证实其临床相关性。结果表明,基本手术动作可在多种场景下被可靠识别,而准确的BSA理解模型能有效支持复杂应用,加速外科智能的实现。
原文摘要 · Abstract (English)
Artificial intelligence, imaging, and large language models have the potential to transform surgical practice, training, and automation. Understanding and modeling of basic surgical actions (BSA), the fundamental unit of operation in any surgery, is important to drive the evolution of this field. In this paper, we present a BSA dataset comprising 10 basic actions across 6 surgical specialties with over 11,000 video clips, which is the largest to date. Based on the BSA dataset, we developed a new foundation model that conducts general-purpose recognition of basic actions. Our approach demonstrates robust cross-specialist performance in experiments validated on datasets from different procedural types and various body parts. Furthermore, we demonstrate downstream applications enabled by the BAS foundation model through surgical skill assessment in prostatectomy using domain-specific knowledge, and action planning in cholecystectomy and nephrectomy using large vision-language models. Multinational surgeons' evaluation of the language model's output of the action planning explainable texts demonstrated clinical relevance. These findings indicate that basic surgical actions can be robustly recognized across scenarios, and an accurate BSA understanding model can essentially facilitate complex applications and speed up the realization of surgical superintelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。