arXiv:2501.15326cs.CV2025-01ICLR被引 5

用弱监督数据训练的手术物体识别模型,零样本识别能力更强。

Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data

  • 从无标注手术视频自动生成图文对,减少人工标注
  • 在四个基准上零样本识别准确率提升2.9到10.6个点
  • 适合需要快速部署手术识别系统的医疗团队

我们提出RASO,一种用于识别任意手术物体的基础模型,具备在多种手术场景和物体类别中对图像与视频进行鲁棒的开放集识别能力。RASO采用新型弱监督学习框架,可从大规模未标注手术讲座视频中自动构建标签-图像-文本对,显著降低人工标注需求。其可扩展的数据生成管道收集了2,200例手术,生成了涵盖2,066个独特标签的360万条标签注释。实验表明,RASO在四个标准手术基准上实现了2.9、4.5、10.6和7.2的mAP提升(零样本设置),并在有监督手术动作识别任务中超越现有最佳模型。代码、模型与演示已公开于https://ntlm1686.github.io/raso。

原文摘要 · Abstract (English)

We present RASO, a foundation model designed to Recognize Any Surgical Object, offering robust open-set recognition capabilities across a broad range of surgical procedures and object classes, in both surgical images and videos. RASO leverages a novel weakly-supervised learning framework that generates tag-image-text pairs automatically from large-scale unannotated surgical lecture videos, significantly reducing the need for manual annotations. Our scalable data generation pipeline gathers 2,200 surgical procedures and produces 3.6 million tag annotations across 2,066 unique surgical tags. Our experiments show that RASO achieves improvements of 2.9 mAP, 4.5 mAP, 10.6 mAP, and 7.2 mAP on four standard surgical benchmarks, respectively, in zero-shot settings, and surpasses state-of-the-art models in supervised surgical action recognition tasks. Code, model, and demo are available at https://ntlm1686.github.io/raso.

手术识别弱监督基础模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。