手术视觉从识别迈向推理,图结构助力理解复杂手术场景。
Decoding the Surgical Scene: A Scoping Review of Scene Graphs in Surgery
- 用图结构建模手术中物体间关系,实现动态环境理解。
- 81%研究基于真实内窥镜视频,但缺乏临床验证。
- 提出三重评估框架,推动图模型走向临床落地。
随着手术人工智能从像素级检测转向复杂推理,场景图(Scene Graphs, SGs)提供了结构化、关系化的表示能力,有助于解析动态手术环境。本研究遵循PRISMA-ScR指南,系统梳理了52项手术领域场景图研究,分析其应用与方法演变。结果表明,该领域发展迅速,但存在显著的‘数据鸿沟’:内部视角研究(如内窥镜视频中的三元组识别)占81%,且几乎全部使用真实2D视频;而外部视角手术室建模则高度依赖模拟数据。方法上,已从基础图神经网络转向专用基础模型与生成式AI,二者合计占2025年研究的约50%。尤为重要的是,场景图正从简单描述工具演变为关键的‘神经符号护栏’,为日益自主的手术基础模型提供可验证的中间表示,防止幻觉。然而,重大转化差距依然存在:所有被评研究均未开展前瞻性临床验证。我们建议超越传统计算机视觉指标,引入‘验证三重奏’——语义查询成功率、延迟感知准确率与安全关键召回率,作为推动图基手术AI进入临床实践的必要评估框架。
原文摘要 · Abstract (English)
As surgical AI transitions from pixel-level detection to complex reasoning, Scene Graphs (SGs) offer the structured, relational representations necessary to decode dynamic surgical environments. This PRISMA-ScR-guided scoping review systematically maps the evolving landscape of SG research in surgery, analyzing 52 primary studies to chart applications and methodological shifts. Our analysis reveals rapid growth, yet uncovers a critical 'data divide': internal-view research (e.g., triplet recognition from endoscopic video) accounts for 81% of studies and almost exclusively uses real-world 2D video, while external-view operating room modeling relies heavily on simulated data. Methodologically, we identify a decisive shift from foundational graph neural networks to specialized foundation models and generative AI, which together now account for approximately 50% of research in 2025. Crucially, our synthesis suggests that Scene Graphs are evolving from simple descriptors into essential 'neuro-symbolic guardrails', providing the structured, verifiable intermediate representation needed to prevent hallucinations in increasingly autonomous Surgical Foundation Models. Despite this promise, a major translational gap remains: none of the reviewed studies have proceeded to prospective clinical validation. We conclude that bridging this gap requires moving beyond standard computer vision metrics; we therefore propose the 'Validation Trinity' -- prioritizing Semantic Query Success, Latency-Aware Accuracy, and Safety-Critical Recall -- as the necessary evaluation framework to bring graph-based surgical AI into clinical practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。