arXiv:2603.09466cs.CV2026-03被引 1

用拓扑结构统一建模手术室多模态关系,提升安全推理能力

TopoOR: A Unified Topological Scene Representation for the Operating Room

  • 将手术室实体关系升维为高阶拓扑单元,保留成对与群体关系
  • 在无菌违规检测、机器人阶段预测等任务上超越传统图模型与大模型基线
  • 专为医疗安全场景设计,适合需精准多模态融合的智能手术系统

手术场景图将手术室复杂性抽象为实体及其关系结构,但现有方法受限于严格二元关系。依赖成对消息传递或分词序列的框架会扁平化本体几何结构,导致关系丢失。本文提出TopoOR,一种将多模态手术室建模为高阶拓扑结构的新范式,天然保留成对与组间关系。通过将实体间交互提升至高阶拓扑胞腔,该表示原生建模手术室中复杂的动态与多模态特性。此拓扑表示涵盖传统场景图,表达能力更强。我们还提出高阶注意力机制,显式保持流形结构与模态特异性特征在层级关系注意力中的完整性。从而避免将3D几何、音频和机器人运动学合并为单一潜在表示,相比现有方法更精确地保留安全关键推理所需的多模态结构。大量实验表明,该方法在无菌违规检测、机器人阶段预测和下一步动作预测任务上均优于传统图模型与基于大语言模型的基线。

原文摘要 · Abstract (English)

Surgical Scene Graphs abstract the complexity of surgical operating rooms (OR) into a structure of entities and their relations, but existing paradigms suffer from strictly dyadic structural limitations. Frameworks that predominantly rely on pairwise message passing or tokenized sequences flatten the manifold geometry inherent to relational structures and lose structure in the process. We introduce TopoOR, a new paradigm that models multimodal operating rooms as a higher-order structure, innately preserving pairwise and group relationships. By lifting interactions between entities into higher-order topological cells, TopoOR natively models complex dynamics and multimodality present in the OR. This topological representation subsumes traditional scene graphs, thereby offering strictly greater expressivity. We also propose a higher-order attention mechanism that explicitly preserves manifold structure and modality-specific features throughout hierarchical relational attention. In this way, we circumvent combining 3D geometry, audio, and robot kinematics into a single joint latent representation, preserving the precise multimodal structure required for safety-critical reasoning, unlike existing methods. Extensive experiments demonstrate that our approach outperforms traditional graph and LLM-based baselines across sterility breach detection, robot phase prediction, and next-action anticipation

拓扑学习手术辅助多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。