用视频自动生成手术流程图,让外科医生几分钟内完成手术分析。
A vision-language model and platform for temporally mapping surgery from video
- 基于65万+手术视频训练视觉语言模型,覆盖8个外科领域。
- 在27,000个视频上超越现有最佳模型,分析更全面且速度快。
- 提供网页平台,医生可直接上传视频自动分析,推动临床落地。
手术映射是制定操作指南和实现自主机器人手术的基础。近年来人工智能在从视频中解析外科医生行为方面取得进展,但现有模型范围有限,仅能捕捉单一术式中的部分行为,且难以转化应用,因缺乏对临床医生的可及性。本文介绍Halsted——一个在涵盖超过65万段视频、覆盖8个外科领域的迭代自标注框架构建的Halsted外科图谱(HSA)上训练的视觉语言模型。为便于基准测试,我们公开发布HSA-27k子集。Halsted在手术行为映射上超越先前最先进模型,兼具更高全面性和计算效率。为弥合外科AI长期存在的转化鸿沟,我们开发了Halsted网页平台(https://halstedhealth.ai/),使全球外科医生能几分钟内自动映射自身手术过程。通过标准化非结构化手术视频数据并直接向医生开放能力,本工作使外科AI更接近临床部署,助力迈向自主机器人手术。
原文摘要 · Abstract (English)
Mapping surgery is fundamental to developing operative guidelines and enabling autonomous robotic surgery. Recent advances in artificial intelligence (AI) have shown promise in mapping the behaviour of surgeons from videos, yet current models remain narrow in scope, capturing limited behavioural components within single procedures, and offer limited translational value, as they remain inaccessible to practising surgeons. Here we introduce Halsted, a vision-language model trained on the Halsted Surgical Atlas (HSA), one of the most comprehensive annotated video libraries grown through an iterative self-labelling framework and encompassing over 650,000 videos across eight surgical specialties. To facilitate benchmarking, we publicly release HSA-27k, a subset of the Halsted Surgical Atlas. Halsted surpasses previous state-of-the-art models in mapping surgical activity while offering greater comprehensiveness and computational efficiency. To bridge the longstanding translational gap of surgical AI, we develop the Halsted web platform (https://halstedhealth.ai/) to provide surgeons anywhere in the world with the previously-unavailable capability of automatically mapping their own procedures within minutes. By standardizing unstructured surgical video data and making these capabilities directly accessible to surgeons, our work brings surgical AI closer to clinical deployment and helps pave the way toward autonomous robotic surgery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。