arXiv:2509.00357cs.CVcs.AI2025-09被引 5

SurgLLM提升手术视频理解,兼顾空间细节与时间顺序。

SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding

  • 用器械中心掩码重建增强空间感知能力
  • 多模态嵌入交错设计提升时间推理效果
  • 动态集成模块适配多种手术任务需求

手术视频理解对计算机辅助手术(CAS)系统至关重要。现有研究仍存在视觉内容感知不足和时间意识薄弱两大瓶颈,限制了通用化CAS解决方案的发展。本文提出SurgLLM框架,一种面向多功能手术视频理解任务的大型多模态模型,具备增强的空间聚焦与时间感知能力。为提升空间感知,我们设计了基于器械中心掩码视频重建(MV-Recon)与多模态对齐的手术上下文感知预训练(Surg-Pretrain)。为融入手术时间知识,提出时间感知多模态微调(TM-Tuning),通过交错多模态嵌入增强时间推理。此外,为避免不同任务间的冲突,设计了手术任务动态集成模块,可高效筛选最优可学习参数。在多种任务上进行的实验,包括生成描述、通用视觉问答(VQA)和时间相关视觉问答(Temporal VQA),均显著优于现有方法,验证了SurgLLM在多功能手术视频理解中的有效性。代码已开源:https://github.com/franciszchen/SurgLLM。

原文摘要 · Abstract (English)

Surgical video understanding is crucial for facilitating Computer-Assisted Surgery (CAS) systems. Despite significant progress in existing studies, two major limitations persist, including inadequate visual content perception and insufficient temporal awareness in surgical videos, and hinder the development of versatile CAS solutions. In this work, we propose the SurgLLM framework, an effective large multimodal model tailored for versatile surgical video understanding tasks with enhanced spatial focus and temporal awareness. Specifically, to empower the spatial focus of surgical videos, we first devise Surgical Context-aware Multimodal Pretraining (Surg-Pretrain) for the video encoder of SurgLLM, by performing instrument-centric Masked Video Reconstruction (MV-Recon) and subsequent multimodal alignment. To incorporate surgical temporal knowledge into SurgLLM, we further propose Temporal-aware Multimodal Tuning (TM-Tuning) to enhance temporal reasoning with interleaved multimodal embeddings. Moreover, to accommodate various understanding tasks of surgical videos without conflicts, we devise a Surgical Task Dynamic Ensemble to efficiently triage a query with optimal learnable parameters in our SurgLLM. Extensive experiments performed on diverse surgical video understanding tasks, including captioning, general VQA, and temporal VQA, demonstrate significant improvements over the state-of-the-art approaches, validating the effectiveness of our SurgLLM in versatile surgical video understanding. The source code is available at https://github.com/franciszchen/SurgLLM.

手术视频多模态时间感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。