用少量标注数据让大模型自动重提示,实现手术视频的连贯分割。
SASVi -- Segment Any Surgical Video
- 基于帧级Mask R-CNN设计自适应重提示机制。
- 相比传统方法,时间一致性提升至少1.5%。
- 适合标注稀缺的医疗视频分割任务。
目的:基础模型虽在大量公开数据上训练,但应用于视觉差异大的目标领域(如手术视频)时,常需额外微调或重提示机制;且缺乏领域知识,难以建模特定语义,导致物体离开或新物体进入场景时无法有效泛化。方法:提出SASVi,一种基于帧级Mask R-CNN Overseer模型的新型重提示机制,仅需少量稀缺标注即可训练。该模型在场景变化时自动触发SAM2的重提示,实现手术视频的时序平滑、完整分割。结果:与类似提示技术及忽略时序信息的帧级分割相比,本方法显著提升时间一致性,至少提高1.5%。我们在三个胆囊切除术和白内障手术数据集上定量与定性验证了SAM2的成功部署。结论:SASVi可作为手术视频连贯分割的新基准,利用稀缺标注生成完整视频标注,并公开这些标注,为未来手术数据科学模型的发展提供丰富数据支持。
原文摘要 · Abstract (English)
Purpose: Foundation models, trained on multitudes of public datasets, often require additional fine-tuning or re-prompting mechanisms to be applied to visually distinct target domains such as surgical videos. Further, without domain knowledge, they cannot model the specific semantics of the target domain. Hence, when applied to surgical video segmentation, they fail to generalise to sections where previously tracked objects leave the scene or new objects enter. Methods: We propose SASVi, a novel re-prompting mechanism based on a frame-wise Mask R-CNN Overseer model, which is trained on a minimal amount of scarcely available annotations for the target domain. This model automatically re-prompts the foundation model SAM2 when the scene constellation changes, allowing for temporally smooth and complete segmentation of full surgical videos. Results: Re-prompting based on our Overseer model significantly improves the temporal consistency of surgical video segmentation compared to similar prompting techniques and especially frame-wise segmentation, which neglects temporal information, by at least 1.5%. Our proposed approach allows us to successfully deploy SAM2 to surgical videos, which we quantitatively and qualitatively demonstrate for three different cholecystectomy and cataract surgery datasets. Conclusion: SASVi can serve as a new baseline for smooth and temporally consistent segmentation of surgical videos with scarcely available annotation data. Our method allows us to leverage scarce annotations and obtain complete annotations for full videos of the large-scale counterpart datasets. We make those annotations publicly available, providing extensive annotation data for the future development of surgical data science models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。