融合视觉与运动数据,提升手术流程识别在干扰下的稳定性。
Multimodal Graph Representation Learning for Robust Surgical Workflow Recognition with Adversarial Feature Disentanglement
- 构建多模态图网络,分离视觉与动作特征以增强表达
- 通过对抗训练对齐跨模态特征,减少数据扰动影响
- 适合需要高鲁棒性的智能手术系统开发人员
手术流程识别对自动化操作、辅助决策和新手培训至关重要,可提升患者安全与流程标准化。然而,出血、烟雾等导致的视觉遮挡或数据存储传输问题会引发性能下降。为此,本文提出一种基于图结构的多模态鲁棒方法,融合视觉与动作数据以提升准确性和可靠性。视觉数据捕捉动态场景,动作数据提供精确运动信息,弥补视觉在恶劣条件下的不足。我们设计了带对抗特征解耦的多模态图表示网络(GRAD),通过图消息传递建模视觉与动作嵌入间的复杂关系;提出视觉-动作对抗框架,利用对抗训练缩小模态间差异,提升特征一致性;并引入上下文校准解码器,融合时序与上下文先验,增强对领域偏移和数据损坏的鲁棒性。大量对比与消融实验验证了模型及模块的有效性。鲁棒性测试表明,该方法在数据存储与传输中仍保持良好稳定性。
原文摘要 · Abstract (English)
Surgical workflow recognition is vital for automating tasks, supporting decision-making, and training novice surgeons, ultimately improving patient safety and standardizing procedures. However, data corruption can lead to performance degradation due to issues like occlusion from bleeding or smoke in surgical scenes and problems with data storage and transmission. In this case, we explore a robust graph-based multimodal approach to integrating vision and kinematic data to enhance accuracy and reliability. Vision data captures dynamic surgical scenes, while kinematic data provides precise movement information, overcoming limitations of visual recognition under adverse conditions. We propose a multimodal Graph Representation network with Adversarial feature Disentanglement (GRAD) for robust surgical workflow recognition in challenging scenarios with domain shifts or corrupted data. Specifically, we introduce a Multimodal Disentanglement Graph Network that captures fine-grained visual information while explicitly modeling the complex relationships between vision and kinematic embeddings through graph-based message modeling. To align feature spaces across modalities, we propose a Vision-Kinematic Adversarial framework that leverages adversarial training to reduce modality gaps and improve feature consistency. Furthermore, we design a Contextual Calibrated Decoder, incorporating temporal and contextual priors to enhance robustness against domain shifts and corrupted data. Extensive comparative and ablation experiments demonstrate the effectiveness of our model and proposed modules. Moreover, our robustness experiments show that our method effectively handles data corruption during storage and transmission, exhibiting excellent stability and robustness. Our approach aims to advance automated surgical workflow recognition, addressing the complexities and dynamism inherent in surgical procedures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。