arXiv:2605.13530cs.CVcs.AI2026-05

用统一模型同时理解手术步骤、器械动作和视觉定位,提升临床辅助准确性。

Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs

论文配图:Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
图 1 · 摘自论文原文
  • 基于多模态大模型,联合建模手术阶段、器械-动作-目标三元组与像素级定位
  • 在CholecT45-Scene数据集上三元组识别准确率提升至46.0%(原40.7%)
  • 适合需要精准手术理解的智能辅助系统研发者

手术场景理解是计算机辅助干预的核心。尽管近期在手术图像分割方面取得进展,但真实临床应用需同时具备过程上下文、语义推理与精确视觉定位的综合理解能力。现有方法通常孤立处理各组件,导致表征碎片化、语义不一致。为此,我们提出SurgMLLM,一个统一的手术场景理解框架,通过单个模型融合高层推理与底层视觉定位。给定手术视频,SurgMLLM微调多模态大语言模型(MLLM),支持结构化可解释推理,联合建模手术阶段、器械-动作-目标(IVT)三元组及三元组-实体分割标记。这些标记经时间聚合后作为提示输入分割网络,实现三元组中器械与目标的像素级定位。整个框架采用端到端训练,以语言推理监督与视觉定位损失联合优化,促进跨任务协同学习与临床一致的场景表征。为支持统一评估,我们引入CholecT45-Scene,扩展原始数据集,提供64,299帧像素级掩码标注,与现有三元组标签对齐。大量实验表明,SurgMLLM显著提升手术场景理解性能,三元组识别主指标AP_IVT从40.7%提升至46.0%,并在阶段识别与分割任务中持续优于现有方法。结果验证了统一推理-定位机制在实现可靠、上下文感知手术辅助中的有效性。

原文摘要 · Abstract (English)

Surgical scene understanding is a cornerstone of computer-assisted intervention. While recent advances, particularly in surgical image segmentation, have driven progress, real-world clinical applications require a more holistic understanding that jointly captures procedural context, semantic reasoning, and precise visual grounding. However, existing approaches typically address these components in isolation, leading to fragmented representations and limited semantic consistency. To address this limitation, we propose SurgMLLM, a unified surgical scene understanding framework that bridges high-level reasoning and low-level visual grounding within a single model. Given surgical videos, SurgMLLM fine-tunes a multimodal large language model (MLLM) to support structured interpretability reasoning, which is used to jointly model phases, instrument-verb-target (IVT) triplets, and triplet-entity segmentation tokens. These tokens are then temporally aggregated and serve as prompts for a segmentation network, enabling accurate pixel-wise grounding of triplet instruments and targets. The entire framework is trained end-to-end with a unified objective that couples language-based reasoning supervision with visual grounding losses, promoting coherent cross-task learning and clinically consistent scene representations. To facilitate unified evaluation, we introduce CholecT45-Scene, extending CholecT45 dataset with 64,299 frames of pixel-level mask annotations for instruments and targets, aligned with existing triplet labels. Extensive experiments show that SurgMLLM significantly advances surgical scene understanding, improving the primary triplet recognition metric AP_IVT from 40.7% to 46.0% and consistently outperforming prior methods in phase recognition and segmentation. These results highlight the effectiveness of unified reasoning-and-grounding for reliable, context-aware surgical assistance.

手术理解多模态模型视觉定位三元组识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。