arXiv:2510.16371cs.CVcs.AI2025-10被引 2

构建3000段白内障手术视频数据集,支持多任务深度学习研究。

Cataract-LMM Large-Scale Multi-Source Multi-Task Benchmark for Deep Learning in Surgical Video Analysis

  • 采集两中心3000段手术视频,涵盖不同医生水平。
  • 提供四层标注:手术阶段、器械结构分割、交互追踪、技能评分。
  • 可训练通用多任务模型,适用于手术分析与能力评估。

计算机辅助手术研究需要大规模、深度标注的视频数据集以捕捉临床与技术变异性。现有白内障手术资源缺乏多样性和标注深度,难以训练具备泛化能力的深度学习模型。为此,我们构建了一个包含3000段超声乳化白内障手术视频的数据集,来自两个手术中心,涵盖不同经验水平的外科医生。该数据集提供四层标注:时间上的手术阶段、器械与解剖结构的实例分割、器械-组织交互追踪,以及基于ICO-OSCAR和GRASIS胜任力量表的定量技能评分。我们通过在四个任务上基准测试深度学习模型,验证了数据集的技术价值:工作流识别、场景分割、器械-组织交互追踪和自动化技能评估。此外,我们建立了一个领域自适应基线,即在一个中心训练,在保留中心评估相位识别与实例分割性能。多源采集、多层标注及配对的技能-运动学标签,推动了可泛化的多任务模型在手术流程分析、场景理解与基于能力的培训研究中的发展。

原文摘要 · Abstract (English)

Computer-assisted surgery research requires large, deeply annotated video datasets that capture clinical and technical variability. Existing cataract surgery resources lack the diversity and annotation depth required to train generalizable deep-learning models. To address this gap, we present a dataset of 3,000 phacoemulsification cataract surgery videos acquired at two surgical centers from surgeons with varying expertise. The dataset provides four annotation layers: temporal surgical phases, instance segmentation of instruments and anatomical structures, instrument-tissue interaction tracking, and quantitative skill scores based on competency rubrics adapted from ICO-OSCAR and GRASIS. We demonstrate the technical utility of the dataset through benchmarking deep learning models across four tasks: workflow recognition, scene segmentation, instrument-tissue interaction tracking, and automated skill assessment. Furthermore, we establish a domain-adaptation baseline for phase recognition and instance segmentation by training on one surgical center and evaluating on a held-out center. Ultimately, these multi-source acquisitions, multi-layer annotations, and paired skill-kinematic labels facilitate the development of generalizable multi-task models for surgical workflow analysis, scene understanding, and competency-based training research.

手术视频多任务数据集技能评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。