arXiv:2507.05007cs.CVcs.AI2025-07

用文本提示提升手术视频安全评估的多标签识别精度

Multi-modal Representations for Fine-grained Multi-label Critical View of Safety Recognition

  • 用正负文本提示对齐图像与描述,实现多标签细粒度分类
  • 在Endoscapes-CVS201上达57.6 mAP,比纯图像模型高6点
  • 适合需轻量标注的外科视觉分析场景

安全腹腔镜胆囊切除术的关键视图(CVS)对手术安全至关重要,但其评估仍具挑战性,即使对专家而言亦然。传统模型依赖昂贵且耗时的空间标注进行纯视觉训练。本文探索如何利用文本作为训练与推理工具,在多模态外科基础模型中自动化CVS识别。不同于多数用于多分类的多模态模型,CVS识别需多标签框架。零样本评估显示现有模型在此任务上表现显著不足。为此,我们提出CVS-AdaptNet,一种通过正负提示对齐图像嵌入与每项CVS标准文本描述的多标签适配策略,增强细粒度二分类性能。在Endoscapes-CVS201数据集上,基于PeskaVLP模型,CVS-AdaptNet达到57.6 mAP,优于仅图像的ResNet50基线(51.5 mAP)6个百分点。结果表明,结合文本提示的多模态多标签框架可显著提升CVS识别效果。我们还提出专用文本推理方法,辅助分析图像-文本对齐。尽管尚未超越基于空间标注的先进方法,但该工作展示了通用模型适配专业外科任务的潜力。

原文摘要 · Abstract (English)

The Critical View of Safety (CVS) is crucial for safe laparoscopic cholecystectomy, yet assessing CVS criteria remains a complex and challenging task, even for experts. Traditional models for CVS recognition depend on vision-only models learning with costly, labor-intensive spatial annotations. This study investigates how text can be harnessed as a powerful tool for both training and inference in multi-modal surgical foundation models to automate CVS recognition. Unlike many existing multi-modal models, which are primarily adapted for multi-class classification, CVS recognition requires a multi-label framework. Zero-shot evaluation of existing multi-modal surgical models shows a significant performance gap for this task. To address this, we propose CVS-AdaptNet, a multi-label adaptation strategy that enhances fine-grained, binary classification across multiple labels by aligning image embeddings with textual descriptions of each CVS criterion using positive and negative prompts. By adapting PeskaVLP, a state-of-the-art surgical foundation model, on the Endoscapes-CVS201 dataset, CVS-AdaptNet achieves 57.6 mAP, improving over the ResNet50 image-only baseline (51.5 mAP) by 6 points. Our results show that CVS-AdaptNet's multi-label, multi-modal framework, enhanced by textual prompts, boosts CVS recognition over image-only methods. We also propose text-specific inference methods, that helps in analysing the image-text alignment. While further work is needed to match state-of-the-art spatial annotation-based methods, this approach highlights the potential of adapting generalist models to specialized surgical tasks. Code: https://github.com/CAMMA-public/CVS-AdaptNet

多模态手术分析细粒度识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。