arXiv:2608.31065cs.CV2026-08

用显微镜视频桥接手术叙述与OCT,实现无需完全同步数据的多模态手术阶段识别。

Multimodal Shared Latent Representation of Narration, Microscope and iOCT Images for Phase Recognition in Vitreoretinal Surgery

论文配图:Multimodal Shared Latent Representation of Narration, Microscope and iOCT Images for Phase Recognition in Vitreoretinal Surgery
图 1 · 摘自论文原文
  • 以显微镜视图为共同锚点,融合手术叙述与iOCT信息
  • 在真实手术中实现宏观阶段识别F1提升至0.53(原0.38)
  • 适用于缺乏同步数据的微创手术智能辅助场景

手术阶段识别对玻璃体视网膜手术中的上下文感知辅助至关重要,但同步多模态术中数据(尤其是显微镜视图与术中OCT)稀缺,限制了模拟外科医生自然多模态整合方法的发展。相比之下,手术叙述在线资源丰富,提供丰富的语义监督。以往工作多聚焦于成对对比学习(如iOCT-显微镜或显微镜-叙述),对三者联合建模研究甚少。本文提出一种框架,利用显微镜视图作为共享锚点,连接手术叙述与术中OCT(iOCT),无需完整同步的三模态数据集,结合真实显微镜-叙述视频与合成的同步显微镜视频与工具对齐iOCT配对数据集。对比对齐将结构先验从合成域迁移到无iOCT的真实视频,双头MS-TCN++融合生成嵌入以实现宏观与微观阶段联合预测。在真实玻璃体视网膜手术上评估,该框架使宏观阶段识别的平均F1从零样本基线0.38提升至0.53,并探索性地估计出真实显微镜视频中不可见的精细器械-组织测量;这些微观阶段估计在合成数据上定量验证,在真实手术中仅定性展示。据我们所知,这是首个将显微镜视图、iOCT B-scan与手术叙述统一于共享潜在空间的手术阶段识别工作。

原文摘要 · Abstract (English)

Surgical phase recognition is key to context-aware computer-assisted feedback in vitreoretinal procedures, yet the scarcity of synchronized multimodal intraoperative data, particularly microscope views and intraoperative OCT, limits approaches that aim to replicate the multimodal integration surgeons perform naturally. Surgical narration, by contrast, is abundantly available online and offers rich semantic supervision. Prior work has mainly explored pairwise contrastive learning (e.g., intraoperative OCT-microscope or microscope-narration), leaving the joint modeling of all three modalities largely unexplored. We introduce a framework that uses microscope views as a shared anchor to bridge surgical narrations and intraoperative OCT (iOCT) without requiring a fully synchronized tri-modal dataset, leveraging real microscope-narration videos and a synthetic dataset of synchronized microscope video and tool-aligned iOCT pairs. Contrastive alignment transfers structural priors from the synthetic domain to real videos lacking iOCT, and a dual-head MS-TCN++ integrates the resulting embeddings for joint macro- and micro-phase prediction. Evaluated on real vitreoretinal surgeries, our framework improves macro-phase recognition over a zero-shot baseline (mean F1 0.38 to 0.53) and provides an exploratory route to estimating fine-grained instrument-tissue measurements that are not directly observable in real microscope video alone; these micro-phase estimates are validated quantitatively on synthetic data and shown only qualitatively on real surgery. To our knowledge, this is the first work to unify microscope view, iOCT B-scans, and surgical narrations in a shared latent space for surgical phase recognition.

手术阶段识别多模态学习OCT显微镜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。