arXiv:2603.05732cs.CV2026-03

用视觉语言对齐自动生成手术流程时间线,省去人工标注。

From Phase Grounding to Intelligent Surgical Narratives

  • 基于CLIP的多模态框架对齐手术视频与动作描述文本。
  • 可自动预测视频帧中的操作手势与阶段,构建结构化时间线。
  • 适合需高效生成手术报告的临床研究与智能手术系统开发。

手术视频时间线在工具辅助手术中至关重要,有助于医生快速定位关键步骤。现有方法依赖术后人工填写报告(常模糊)或手动标注视频(耗时),我们提出介于两者之间的自动化方案:直接从手术视频生成时间线与叙事。采用基于CLIP的多模态框架,将手术视频帧与对应动作描述文本映射至共享嵌入空间。使用视觉编码器提取视频帧特征,文本编码器嵌入动作语句,通过微调增强视频动作与文本标记间的对齐。训练完成后,模型可预测每帧的动作与阶段,实现结构化时间线构建。该方法利用预训练多模态表示,连接视觉动作与文本叙述,显著减少医生对视频的逐帧审查与标注需求。

原文摘要 · Abstract (English)

Video surgery timelines are an important part of tool-assisted surgeries, as they allow surgeons to quickly focus on key parts of the procedure. Current methods involve the surgeon filling out a post-operation (OP) report, which is often vague, or manually annotating the surgical videos, which is highly time-consuming. Our proposed method sits between these two extremes: we aim to automatically create a surgical timeline and narrative directly from the surgical video. To achieve this, we employ a CLIP-based multi-modal framework that aligns surgical video frames with textual gesture descriptions. Specifically, we use the CLIP visual encoder to extract representations from surgical video frames and the text encoder to embed the corresponding gesture sentences into a shared embedding space. We then fine-tune the model to improve the alignment between video gestures and textual tokens. Once trained, the model predicts gestures and phases for video frames, enabling the construction of a structured surgical timeline. This approach leverages pretrained multi-modal representations to bridge visual gestures and textual narratives, reducing the need for manual video review and annotation by surgeons.

手术分析多模态时间线生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。