arXiv:2510.12904cs.CV2025-10

通过状态变化学习预测内窥镜手术未来事件,提升手术安全与效率。

State-Change Learning for Prediction of Future Events in Endoscopic Videos

  • 将未来预测转为状态转移分类,避免直接预测复杂视频帧。
  • 在四个数据集上实现跨任务、跨术式的显著性能提升。
  • 适合关注手术智能预警与自动化支持的研究者与临床工程师。

手术未来预测依赖于对实时手术视频的AI分析,对手术室安全与效率至关重要。它能提前提供未来事件的时间、内容及风险提示,助力资源调度、器械准备和并发症(如出血、胆管损伤)预警。然而,现有研究多聚焦于当前动作理解,缺乏统一方法覆盖短期(动作三元组、事件)与长期(剩余手术时长、阶段与步骤转换)预测任务。现有方法依赖粗粒度监督,对细粒度动作三元组与步骤关注不足;仅基于未来特征预测的方法难以跨术式泛化。为此,本文将手术未来预测重构为状态变化学习:不直接预测原始观测,而是分类当前与未来时间步之间的状态转移。提出SurgFUTR,采用教师-学生架构:视频片段经Sinkhorn-Knopp聚类压缩为状态表示;教师从当前与未来片段中学习,学生仅凭当前视频预测未来状态,由动作动力学(ActDyn)模块引导。构建SFPBench基准,包含五项跨越短/长期预测的任务。在四个数据集、三种术式上的实验显示一致改进,跨术式迁移验证了模型泛化能力。

原文摘要 · Abstract (English)

Surgical future prediction, driven by real-time AI analysis of surgical video, is critical for operating room safety and efficiency. It provides actionable insights into upcoming events, their timing, and risks-enabling better resource allocation, timely instrument readiness, and early warnings for complications (e.g., bleeding, bile duct injury). Despite this need, current surgical AI research focuses on understanding what is happening rather than predicting future events. Existing methods target specific tasks in isolation, lacking unified approaches that span both short-term (action triplets, events) and long-term horizons (remaining surgery duration, phase transitions). These methods rely on coarse-grained supervision while fine-grained surgical action triplets and steps remain underexplored. Furthermore, methods based only on future feature prediction struggle to generalize across different surgical contexts and procedures. We address these limits by reframing surgical future prediction as state-change learning. Rather than forecasting raw observations, our approach classifies state transitions between current and future timesteps. We introduce SurgFUTR, implementing this through a teacher-student architecture. Video clips are compressed into state representations via Sinkhorn-Knopp clustering; the teacher network learns from both current and future clips, while the student network predicts future states from current videos alone, guided by our Action Dynamics (ActDyn) module. We establish SFPBench with five prediction tasks spanning short-term (triplets, events) and long-term (remaining surgery duration, phase and step transitions) horizons. Experiments across four datasets and three procedures show consistent improvements. Cross-procedure transfer validates generalizability.

手术预测状态变化视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。