arXiv:2602.05638cs.CV2026-02被引 5

用运动预测替代像素重建,提升手术视频理解能力

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

  • 以潜在运动预测为核心,聚焦手术语义结构而非视觉细节
  • 在17个基准上显著优于现有方法,最高提升14.6%的F1分数
  • 适合医疗AI研究者与手术自动化系统开发者使用

尽管基础模型已推动手术视频分析发展,但当前方法多依赖像素级重建目标,浪费模型能力于烟雾、反光、液体流动等低层视觉细节,而非对理解手术至关重要的语义结构。我们提出SurgMotion,一种视频原生的基础模型,将学习范式从像素级重建转向潜在运动预测。基于视频联合嵌入预测架构(V-JEPA),SurgMotion引入三项关键技术:(1) 运动引导的潜在掩码预测,优先关注语义重要区域;(2) 空时关联自蒸馏,强化关系一致性;(3) 空时特征多样性正则化(SFDR),防止纹理稀疏场景下的表征坍缩。为支持大规模预训练,我们构建了目前最大的手术视频数据集SurgMotion-15M,包含来自13个解剖部位、50个来源的3,658小时视频。在17个基准上的大量实验表明,SurgMotion在手术流程识别上取得显著提升,于EgoSurgery数据集上F1分数提高14.6%,在PitVis上提升10.3%;在动作三元组识别上,CholecT50数据集mAP-IVT达39.54%;同时在技能评估、息肉分割和深度估计任务中表现优异。这些结果确立了SurgMotion作为通用、以运动为导向的手术视频理解新标准。

原文摘要 · Abstract (English)

While foundation models have advanced surgical video analysis, current approaches rely predominantly on pixel-level reconstruction objectives that waste model capacity on low-level visual details, such as smoke, specular reflections, and fluid motion, rather than semantic structures essential for surgical understanding. We present SurgMotion, a video-native foundation model that shifts the learning paradigm from pixel-level reconstruction to latent motion prediction. Built on the Video Joint Embedding Predictive Architecture (V-JEPA), SurgMotion introduces three key technical innovations tailored to surgical videos: (1) motion-guided latent masked prediction to prioritize semantically meaningful regions, (2) spatiotemporal affinity self-distillation to enforce relational consistency, and (3) spatiotemporal feature diversity regularization (SFDR) to prevent representation collapse in texture-sparse surgical scenes. To enable large-scale pretraining, we curate SurgMotion-15M, the largest surgical video dataset to date, comprising 3,658 hours of video from 50 sources across 13 anatomical regions. Extensive experiments across 17 benchmarks demonstrate that SurgMotion significantly outperforms state-of-the-art methods on surgical workflow recognition, achieving 14.6 percent improvement in F1 score on EgoSurgery and 10.3 percent on PitVis; on action triplet recognition with 39.54 percent mAP-IVT on CholecT50; as well as on skill assessment, polyp segmentation, and depth estimation. These results establish SurgMotion as a new standard for universal, motion-oriented surgical video understanding.

手术视频视频理解基础模型运动预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。