arXiv:2608.24541cs.CV2026-08

用分层原型记忆提升SAM在手术器械分割中的鲁棒性

Hierarchical Prototype-Memory Adaptation of SAM for Surgical Instrument Segmentation

论文配图:Hierarchical Prototype-Memory Adaptation of SAM for Surgical Instrument Segmentation
图 1 · 摘自论文原文
  • 构建多尺度冻结原型库,通过轻量适配器融入SAM特征空间
  • 在EndoVis2017/2018上达到当前最优性能,显著提升复杂场景分割精度
  • 适合需要高精度医学图像分割的临床辅助系统开发者

手术器械分割(SIS)是计算机辅助手术的基础,可靠的器械掩码可实现精准场景理解与临床支持。近期通过提示学习将基础模型如分割一切模型(SAM)适配至手术领域取得良好效果,但在复杂术中条件下表现受限于优化机制不足。单纯依赖下游分割损失优化提示或原型,易导致其退化为特定任务参数,难以维持稳定类别记忆,降低对术中变化的鲁棒性。同时,多尺度视觉线索经单一提示路径传输形成瓶颈,阻碍有效尺度匹配耦合。为此,本文提出层级原型记忆适配框架HPMA。HPMA从标注手术场景构建冻结的多尺度视觉原型记忆库,并通过轻量适配器嵌入SAM特征空间,以保留稳定类别证据。为最大化多尺度线索效用,引入尺度匹配耦合机制:全局原型校准类别级提示特征,结构原型引导解码器对象查询,局部原型通过局部对齐目标与高分辨率特征图对齐。在公开数据集EndoVis2017和EndoVis2018上的大量实验表明,该方法性能达到当前最优,超越现有基础模型适配方法。

原文摘要 · Abstract (English)

Surgical instrument segmentation (SIS) is fundamental for computer-assisted surgery, where reliable instrument masks enable precise scene understanding and clinical assistance. Recently, adapting foundation models like the Segment Anything Model (SAM) to the surgical domain via prompt-learning has shown encouraging results. However, the performance of these adapted models under challenging surgical conditions is constrained by suboptimal adaptation mechanisms. Specifically, optimizing prompts or prototypes purely via downstream segmentation loss tends to cause them to degenerate into task-specific parameters rather than serving as persistent, stable category memory, thereby degrading their robustness against complex intraoperative variations. Moreover, routing multi-scale visual cues through a single prompt pathway creates a bottleneck that hinders effective scale-matched coupling. To address these limitations, we propose HPMA, a Hierarchical Prototype-Memory Adaptation framework for SAM. Specifically, HPMA constructs a frozen, multi-scale visual prototype memory bank from annotated surgical scenes and integrates it into SAM's feature space using lightweight adapters to preserve stable category evidence. To maximize the utility of multi-scale cues, we introduce a scale-matched coupling mechanism where global prototypes calibrate class-level prompt features, structural prototypes guide decoder object queries, and local prototypes align high-resolution feature maps through a local alignment objective. Extensive experiments on the public EndoVis2017 and EndoVis2018 datasets demonstrate that our approach achieves state-of-the-art performance, outperforming existing foundation model adaptation methods.

医学图像分割原型记忆SAM适配手术导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。