构建超大内窥镜数据集LEMON并训练新模型,提升手术视觉理解性能。
LEMON: A Large Endoscopic MONocular Dataset and Foundation Model for Perception in Surgical Settings
- 用新型聚合管道收集4000+小时高质量手术视频,覆盖多种术式。
- 新模型LemonFM在6个数据集上多项任务显著超越现有模型,最高提升10.3个百分点。
- 适合研究自主手术机器人、医学影像分析的团队使用,开源可复现。
传统公开手术数据集规模小,通常不足100段视频、30小时内容,导致模型泛化能力差。为解决这一问题,本文构建了名为LEMON的新数据集,通过创新聚合流程从网络来源收集高分辨率视频。LEMON包含超过4000段手术视频,总计938小时(8500万帧)高质量影像,覆盖多种手术类型,规模与多样性均超越现有资源,并引入两项新下游任务。为验证其有效性,我们提出LemonFM,一种基于新型自监督增强知识蒸馏方法在LEMON上预训练的通用模型。LemonFM在四个下游任务和六个数据集上持续优于现有手术基础模型:手术阶段识别(AutoLaparo、M2CAI16、Cholec80的Jaccard分别提升9.5、9.4、8.4个百分点),手术动作识别(CholecT50的mAP提升4.4个百分点),手术器械存在检测(Cholec80和GraSP的mAP分别提升5.3、10.2个百分点),手术语义分割(CholecSeg8k的mDice提升10.3个百分点)。LEMON与LemonFM将为科研与产业界提供基础资源,推动自主手术机器人发展,助力更安全、普惠的全球外科医疗。数据、代码与模型已开源于https://github.com/visurg-ai/LEMON。
原文摘要 · Abstract (English)
Traditional open-access datasets focusing on surgical procedures are often limited by their small size, typically consisting of fewer than 100 videos and less than 30 hours of footage, which leads to poor model generalization. To address this data limitation, a new dataset called LEMON has been compiled using a novel aggregation pipeline that collects high-resolution videos from online sources. Featuring an extensive collection of over 4K surgical videos totaling 938 hours (85 million frames) of high-quality footage across multiple procedure types, LEMON offers a comprehensive resource surpassing existing alternatives in size and scope, including two novel downstream tasks. To demonstrate the effectiveness of this diverse dataset, we introduce LemonFM, a foundation model pretrained on LEMON using a novel self-supervised augmented knowledge distillation approach. LemonFM consistently outperforms existing surgical foundation models across four downstream tasks and six datasets, achieving significant gains in surgical phase recognition (+9.5pp, +9.4pp, and +8.4pp in Jaccard on AutoLaparo, M2CAI16, and Cholec80), surgical action recognition (+4.4pp in mAP on CholecT50), surgical tool presence detection (+5.3pp and +10.2pp in mAP on Cholec80 and GraSP), and surgical semantic segmentation (+10.3pp in mDice on CholecSeg8k). LEMON and LemonFM will serve as foundational resources for the research community and industry, accelerating progress in developing autonomous robotic surgery systems and ultimately contributing to safer and more accessible surgical care worldwide. Dataset, code, and models are publicly available at https://github.com/visurg-ai/LEMON.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。