用动作图生成腹腔镜手术未来视频,提升数据稀缺的外科研究能力。
VISAGE: Video Synthesis using Action Graphs for Surgery
- 基于动作场景图建模手术流程顺序,结合扩散模型生成视频。
- 仅需初始帧和动作图三元组即可预测后续帧,保持时空一致性。
- 适合外科数据增强、手术模拟与机器人辅助系统开发。
外科数据科学(SDS)分析术前、术中和术后患者数据以改善手术结果与技能。然而,外科数据稀缺、异构且复杂,限制了现有机器学习方法的应用。本文提出腹腔镜手术未来视频生成的新任务,可扩充和丰富现有数据,支持仿真、分析及机器人辅助手术等应用。该任务不仅需理解当前手术状态,还需准确预测手术过程的动态与不可预测性。我们提出的VISAGE(VIdeo Synthesis using Action Graphs for Surgery)方法利用动作场景图捕捉腹腔镜手术的时序特征,并通过扩散模型生成时间连贯的视频序列。VISAGE在仅输入单个初始帧和动作图三元组的情况下,预测未来帧。通过引入领域知识的动作图,确保生成视频符合真实腹腔镜手术中的视觉与运动模式。实验结果表明,VISAGE能实现高保真度的腹腔镜手术视频生成,为外科数据科学提供多种应用支持。
原文摘要 · Abstract (English)
Surgical data science (SDS) is a field that analyzes patient data before, during, and after surgery to improve surgical outcomes and skills. However, surgical data is scarce, heterogeneous, and complex, which limits the applicability of existing machine learning methods. In this work, we introduce the novel task of future video generation in laparoscopic surgery. This task can augment and enrich the existing surgical data and enable various applications, such as simulation, analysis, and robot-aided surgery. Ultimately, it involves not only understanding the current state of the operation but also accurately predicting the dynamic and often unpredictable nature of surgical procedures. Our proposed method, VISAGE (VIdeo Synthesis using Action Graphs for Surgery), leverages the power of action scene graphs to capture the sequential nature of laparoscopic procedures and utilizes diffusion models to synthesize temporally coherent video sequences. VISAGE predicts the future frames given only a single initial frame, and the action graph triplets. By incorporating domain-specific knowledge through the action graph, VISAGE ensures the generated videos adhere to the expected visual and motion patterns observed in real laparoscopic procedures. The results of our experiments demonstrate high-fidelity video generation for laparoscopy procedures, which enables various applications in SDS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。