用手术语法结构提升腹腔镜手术意图识别准确率
Leveraging Surgical Activity Grammar for Primary Intention Prediction in Laparoscopy Procedures
- 结合手术语法与视觉检测,构建上下文感知的意图识别框架
- 在基准数据集上优于仅依赖视觉特征的现有方法
- 适合研究手术自动化与智能机器人系统的开发者
外科手术本质上复杂且动态,存在错综复杂的依赖关系和多种执行路径。准确识别关键操作背后的意图(即主意图,PI)对于理解与规划手术过程至关重要。本文提出一种新框架,通过融合自上而下的语法结构与自下而上的视觉线索,提升手术视频中的主意图识别能力。该语法结构基于大量手术流程语料库,提供手术活动的分层视角。一个基于手术活动语法的解析器,利用手术动作检测器获取的腹腔镜图像视觉数据进行处理,从而更精准地解读视觉信息。在基准数据集上的实验结果表明,本方法显著优于仅依赖视觉特征的现有手术动作检测器。研究成果为开发具备更强规划与自动化能力的智能手术机器人系统奠定了良好基础。
原文摘要 · Abstract (English)
Surgical procedures are inherently complex and dynamic, with intricate dependencies and various execution paths. Accurate identification of the intentions behind critical actions, referred to as Primary Intentions (PIs), is crucial to understanding and planning the procedure. This paper presents a novel framework that advances PI recognition in instructional videos by combining top-down grammatical structure with bottom-up visual cues. The grammatical structure is based on a rich corpus of surgical procedures, offering a hierarchical perspective on surgical activities. A grammar parser, utilizing the surgical activity grammar, processes visual data obtained from laparoscopic images through surgical action detectors, ensuring a more precise interpretation of the visual information. Experimental results on the benchmark dataset demonstrate that our method outperforms existing surgical activity detectors that rely solely on visual features. Our research provides a promising foundation for developing advanced robotic surgical systems with enhanced planning and automation capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。