arXiv:2503.02904eess.IVcs.CV2025-03中稿 · the Data Engineeri…被引 14

首个可生成可控手术视频的视觉世界模型,无需动作标注。

Surgical Vision World Model

  • 基于无标签手术视频,自动推断隐含操作动作。
  • 在SurgToolLoc-2022数据集上实现动作可控视频生成。
  • 适合医学仿真、智能手术助手训练场景。

真实且可交互的外科手术模拟有望推动医疗专业培训与自主外科智能体训练等关键应用。在自然视觉领域,世界模型已实现动作控制的数据生成,证明了在大规模真实数据难以获取时,可在交互式仿真环境中训练自主智能体。然而,现有外科领域的研究仍局限于简化计算机模拟,缺乏真实感。此外,现有世界模型多依赖带动作标注的数据,限制其在真实外科数据中的应用,因动作标注成本过高。受近期Genie模型利用无标签游戏视频推断潜在动作并实现动作控制生成的启发,我们提出首个外科视觉世界模型。该模型可生成动作可控的外科数据,其架构设计已在无标签的SurgToolLoc-2022数据集上通过大量实验验证。代码与实现细节见GitHub:https://github.com/bhattarailab/Surgical-Vision-World-Model。

原文摘要 · Abstract (English)

Realistic and interactive surgical simulation has the potential to facilitate crucial applications, such as medical professional training and autonomous surgical agent training. In the natural visual domain, world models have enabled action-controlled data generation, demonstrating the potential to train autonomous agents in interactive simulated environments when large-scale real data acquisition is infeasible. However, such works in the surgical domain have been limited to simplified computer simulations, and lack realism. Furthermore, existing literature in world models has predominantly dealt with action-labeled data, limiting their applicability to real-world surgical data, where obtaining action annotation is prohibitively expensive. Inspired by the recent success of Genie in leveraging unlabeled video game data to infer latent actions and enable action-controlled data generation, we propose the first surgical vision world model. The proposed model can generate action-controllable surgical data and the architecture design is verified with extensive experiments on the unlabeled SurgToolLoc-2022 dataset. Codes and implementation details are available at https://github.com/bhattarailab/Surgical-Vision-World-Model

手术模拟世界模型视频生成无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。