构建含器械-动作-目标与操作手身份的手术场景图,提升智能手术系统理解力
Towards Holistic Surgical Scene Graph
- 提出基于图结构的建模方法,融合器械-动作-目标组合与操作手身份
- 在关键视图安全评估与动作三元组识别任务上显著提升性能
- 公开了首个包含操作手身份标注的手术场景图数据集,适合医疗视觉研究者
手术场景理解对计算机辅助干预系统至关重要,需对包括手术器械、解剖结构及其交互在内的多种元素进行视觉认知。为有效表征手术场景的复杂信息,已有研究探索了基于图的方法来结构化建模手术实体及其关系。先前研究虽证明了图表示手术场景的可行性,但器械-动作-目标的多样化组合以及操作器械的手部身份等关键方面仍缺乏充分建模。为此,本文提出 Endoscapes-SG201 数据集,包含工具-动作-目标组合与手部身份的标注。同时引入 SSG-Com 方法,专门学习并表示这些关键要素。在关键视图安全评估与动作三元组识别等下游任务上的实验表明,整合这些组件显著提升了手术场景理解能力。代码与数据集已开源。
原文摘要 · Abstract (English)
Surgical scene understanding is crucial for computer-assisted intervention systems, requiring visual comprehension of surgical scenes that involves diverse elements such as surgical tools, anatomical structures, and their interactions. To effectively represent the complex information in surgical scenes, graph-based approaches have been explored to structurally model surgical entities and their relationships. Previous surgical scene graph studies have demonstrated the feasibility of representing surgical scenes using graphs. However, certain aspects of surgical scenes-such as diverse combinations of tool-action-target and the identity of the hand operating the tool-remain underexplored in graph-based representations, despite their importance. To incorporate these aspects into graph representations, we propose Endoscapes-SG201 dataset, which includes annotations for tool-action-target combinations and hand identity. We also introduce SSG-Com, a graph-based method designed to learn and represent these critical elements. Through experiments on downstream tasks such as critical view of safety assessment and action triplet recognition, we demonstrated the importance of integrating these essential scene graph components, highlighting their significant contribution to surgical scene understanding. The code and dataset are available at https://github.com/ailab-kyunghee/SSG-Com
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。