arXiv:2506.01980eess.IVcs.AI2025-06

用压缩与熵最大化学习手术视频的紧凑表征,无需标注数据即可提升多种手术辅助任务性能。

Surgical Foundation Model Leveraging Compression and Entropy Maximization for Image-Guided Surgical Assistance

  • 基于柯尔莫戈洛夫复杂度设计自监督框架,通过熵最大化解码器压缩图像保留关键信息。
  • 在无标注手术数据上训练后,在工作流程分类等任务上实现显著性能提升。
  • 模型内部表征能更好解耦图像结构特征,适合医疗视觉基础模型研究者使用。

实时视频理解对微创手术(MIS)中的操作引导至关重要。然而,监督学习需要大量标注数据,而医学领域因标注成本过高导致数据稀缺。尽管自监督方法可缓解此问题,但现有方法往往无法有效捕捉可泛化的结构与物理信息。本文提出压缩-探索(Compress-to-Explore, C2E)框架,利用柯尔莫戈洛夫复杂度从手术视频中学习紧凑且信息丰富的表征。C2E通过熵最大化解码器在压缩图像的同时保留临床相关细节,从而提升编码器性能,且无需标签数据。在大规模未标注手术数据集上训练后,该模型在工作流程分类、器械-组织交互分类、分割及诊断等多样任务中展现出强泛化能力,作为手术视觉基础模型表现优异。进一步研究表明,模型内部紧凑表征能更好解耦图像中不同结构部分的特征。这些结果凸显了自监督学习在提升手术人工智能方面的巨大潜力,有望改善微创手术效果。

原文摘要 · Abstract (English)

Real-time video understanding is critical to guide procedures in minimally invasive surgery (MIS). However, supervised learning approaches require large, annotated datasets that are scarce due to annotation efforts that are prohibitive, e.g., in medical fields. Although self-supervision methods can address such limitations, current self-supervised methods often fail to capture structural and physical information in a form that generalizes across tasks. We propose Compress-to-Explore (C2E), a novel self-supervised framework that leverages Kolmogorov complexity to learn compact, informative representations from surgical videos. C2E uses entropy-maximizing decoders to compress images while preserving clinically relevant details, improving encoder performance without labeled data. Trained on large-scale unlabeled surgical datasets, C2E demonstrates strong generalization across a variety of surgical ML tasks, such as workflow classification, tool-tissue interaction classification, segmentation, and diagnosis tasks, providing improved performance as a surgical visual foundation model. As we further show in the paper, the model's internal compact representation better disentangles features from different structural parts of images. The resulting performance improvements highlight the yet untapped potential of self-supervised learning to enhance surgical AI and improve outcomes in MIS.

自监督学习手术辅助视觉基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。