arXiv:2508.10215eess.IVcs.CV2025-08被引 2

用少量标注数据实现更通用的手术视频理解,提升临床可用性。

Data-Efficient Learning for Generalizable Surgical Video Understanding

  • 提出DIST、SemiVT-Surge等半监督框架,利用无标签视频提升模型性能。
  • 在极低标注数据下达成当前最优结果,动态伪标签有效缓解数据稀缺问题。
  • 发布两个大型多任务手术数据集,推动可复现研究与临床落地。

外科视频分析的进步正推动手术室向智能化、数据驱动的环境转变。计算机辅助系统支持从术前规划到术中引导和术后评估的全流程。然而,由于标注稀缺、时空复杂性以及不同手术和机构间的领域差异,构建鲁棒且泛化的外科视频理解模型仍具挑战。本研究旨在弥合深度学习在科研与临床部署之间的差距。为解决关键任务如手术阶段、操作和事件识别,我们基准测试了主流神经网络架构,并提出新模型与先进模块以提升性能。针对专家标注成本高及跨源领域差异问题,重点降低对标注数据的依赖。我们开发了半监督框架,通过大量无标签视频提升多任务表现。提出的DIST、SemiVT-Surge和ENCORE框架,在仅用少量标注数据的情况下,于多个挑战性外科数据集上达到领先水平,其核心是动态伪标签机制。为促进可复现性与领域进步,我们发布了两个多任务数据集:目前最大的妇科腹腔镜数据集GynSurg和最大白内障手术视频数据集Cataract-1K。本工作为外科视频分析提供了鲁棒、数据高效且可临床扩展的解决方案,奠定了可泛化AI系统应用于外科诊疗与培训的基础。

原文摘要 · Abstract (English)

Advances in surgical video analysis are transforming operating rooms into intelligent, data-driven environments. Computer-assisted systems support full surgical workflow, from preoperative planning to intraoperative guidance and postoperative assessment. However, developing robust and generalizable models for surgical video understanding remains challenging due to (I) annotation scarcity, (II) spatiotemporal complexity, and (III) domain gap across procedures and institutions. This doctoral research aims to bridge the gap between deep learning-based surgical video analysis in research and its real-world clinical deployment. To address the core challenge of recognizing surgical phases, actions, and events, critical for analysis, I benchmarked state-of-the-art neural network architectures to identify the most effective designs for each task. I further improved performance by proposing novel architectures and integrating advanced modules. Given the high cost of expert annotations and the domain gap across surgical video sources, I focused on reducing reliance on labeled data. We developed semi-supervised frameworks that improve model performance across tasks by leveraging large amounts of unlabeled surgical video. We introduced novel semi-supervised frameworks, including DIST, SemiVT-Surge, and ENCORE, that achieved state-of-the-art results on challenging surgical datasets by leveraging minimal labeled data and enhancing model training through dynamic pseudo-labeling. To support reproducibility and advance the field, we released two multi-task datasets: GynSurg, the largest gynecologic laparoscopy dataset, and Cataract-1K, the largest cataract surgery video dataset. Together, this work contributes to robust, data-efficient, and clinically scalable solutions for surgical video analysis, laying the foundation for generalizable AI systems that can meaningfully impact surgical care and training.

手术视频半监督数据效率医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。