arXiv:2603.21566cs.CVcs.AI2026-03

基于SAM2的白内障手术视频分割模型,支持实时高精度分割与高效标注。

CataractSAM-2: A Domain-Adapted Model for Anterior Segment Surgery Segmentation and Scalable Ground-Truth Annotation

  • 基于SAM2改进,适配眼科前段手术视频实时分割
  • 零样本泛化至青光眼手术,跨术式通用性强
  • 结合稀疏提示与视频掩码传播,标注效率提升显著

我们提出CataractSAM-2,是Meta Segment Anything Model 2的领域自适应扩展,专为白内障眼科手术视频实现高精度实时语义分割。该模型处于计算机视觉与医疗机器人交叉领域,可支持机器人辅助及计算机引导手术系统的术中感知。为减轻人工标注负担,我们引入一种交互式标注框架,结合稀疏提示与视频掩码传播技术,显著降低标注时间,推动高质量真值掩码的规模化生成,加速眼科前段手术数据集构建。我们还验证了模型在青光眼小梁切除术上的强零样本泛化能力,证明其跨术式适用性及更广泛外科应用潜力。训练好的模型与标注工具已开源,确立CataractSAM-2作为扩展前段眼科手术数据集、推进实时AI驱动医疗机器人与手术视频理解的基石。

原文摘要 · Abstract (English)

We present CataractSAM-2, a domain-adapted extension of Meta's Segment Anything Model 2, designed for real-time semantic segmentation of cataract ophthalmic surgery videos with high accuracy. Positioned at the intersection of computer vision and medical robotics, CataractSAM-2 enables precise intraoperative perception crucial for robotic-assisted and computer-guided surgical systems. Furthermore, to alleviate the burden of manual labeling, we introduce an interactive annotation framework that combines sparse prompts with video-based mask propagation. This tool significantly reduces annotation time and facilitates the scalable creation of high-quality ground-truth masks, accelerating dataset development for ocular anterior segment surgeries. We also demonstrate the model's strong zero-shot generalization to glaucoma trabeculectomy procedures, confirming its cross-procedural utility and potential for broader surgical applications. The trained model and annotation toolkit are released as open-source resources, establishing CataractSAM-2 as a foundation for expanding anterior ophthalmic surgical datasets and advancing real-time AI-driven solutions in medical robotics, as well as surgical video understanding.

医学图像分割实时分割医疗机器人自动标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。