提出SurgTEMP框架,让AI理解腹腔镜胆囊切除术视频中的时间动态与手术知识。
SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy
- 用查询引导的视觉记忆构建时空双层记忆库,捕捉手术过程变化。
- 在32000个问答对上提升多任务评估性能,尤其在高级认知任务中显著领先。
- 适合医学教育、术中辅助等需要理解复杂手术流程的场景。
外科手术过程高度复杂且风险高,需专家级知识与持续专注力以应对不断变化的术中画面。计算机辅助系统如手术视觉问答(VQA)在教学与术中支持方面具有潜力。现有研究多聚焦静态图像分析,忽视丰富的时序语义。手术视频问答还面临视觉对比度低、强知识依赖性、跨时段分析需求多样、从基础感知到高级评估的层次化特征等挑战。为此,我们提出SurgTEMP,一种多模态大模型框架,包含(i)查询引导的标记选择模块,用于构建层级化视觉记忆(空间与时间记忆库),以及(ii)手术能力进展(SCP)训练方案。二者协同实现对可变长度手术视频的有效建模,保留关键术中线索与时间连贯性,并更好支持多样化下游评估任务。为推动模型发展,我们引入CholeVidQA-32K数据集,包含32,000个开放问答对和3,855段视频片段(约128小时),源自腹腔镜胆囊切除术,按感知、评估、推理三层次组织,涵盖11项任务,从器械/动作/解剖感知到安全视图(CVS)、术中难度、技能水平与不良事件评估。在与先进开源多模态及视频大模型(微调与零样本)的全面对比中,SurgTEMP取得显著性能提升,推动了基于视频的手术VQA新基准。
原文摘要 · Abstract (English)
Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes. Computer-assisted systems such as surgical visual question answering (VQA) offer promises for education and intraoperative support. Current surgical VQA research largely focuses on static frame analysis, overlooking rich temporal semantics. Surgical video question answering is further challenged by low visual contrast, its highly knowledge-driven nature, diverse analytical needs spanning scattered temporal windows, and the hierarchy from basic perception to high-level intraoperative assessment. To address these challenges, we propose SurgTEMP, a multimodal LLM framework featuring (i) a query-guided token selection module that builds hierarchical visual memory (spatial and temporal memory banks) and (ii) a Surgical Competency Progression (SCP) training scheme. Together, they enable effective modeling of variable-length surgical videos while preserving procedure-relevant cues and temporal coherence, and better support diverse downstream assessment tasks. To support model development, we introduce CholeVidQA-32K, a surgical video question answering dataset comprising 32K open-ended QA pairs and 3,855 video segments (approximately 128 h total) from laparoscopic cholecystectomy. The dataset is organized into a three-level hierarchy -- Perception, Assessment, and Reasoning -- spanning 11 tasks from instrument/action/anatomy perception to Critical View of Safety (CVS), intraoperative difficulty, skill proficiency, and adverse event assessment. In comprehensive evaluations against state-of-the-art open-source multimodal and video LLMs (fine-tuned and zero-shot), SurgTEMP achieves substantial performance improvements, advancing the state of video-based surgical VQA. The project page is available at: https://camma-public.github.io/SurgTEMP/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。