综述外科视频生成技术,梳理从图像合成到动态建模的演进路径。
Surgical Video Generation From Diffusion to World Models: A Survey

- 按无条件、条件生成和世界模型三类归纳2024-2026年研究
- 发现像素保真与临床合理性的核心矛盾,量化评估主流方法表现
- 适合智能感知、生成AI与手术数据科学交叉研究者参考
外科视频数据是术中感知、手术流程理解与机器人决策模型的主要训练资源,但临床数据获取受限于隐私、成本及类别不平衡。外科视频生成成为缓解数据稀缺的关键方法,为手术模拟、训练与机器人策略学习奠定基础。该领域发展迅速却缺乏清晰概念框架。本文将2024–2026年文献归纳为三类:无条件生成、条件生成与世界模型生成,揭示任务定义正从生成视觉上逼真的帧,转向建模手术场景的因果动态。我们分析像素级保真与临床合理性间的持续差距,指出泛化性、物理真实性、可控性与可解释性为关键瓶颈。进一步总结代表性方法在公开数据集上的实验结果,提供量化参考。本综述系统梳理当前进展与开放挑战,为智能感知、多模态融合、生成式AI与外科数据科学交叉领域的研究者提供指导。
原文摘要 · Abstract (English)
Surgical video data provides the primary training resource for models of intraoperative perception, surgical workflow understanding, and robotic decision-making. However, clinical data acquisition remains constrained by privacy, cost, and class imbalance. Surgical video generation has emerged as a transformative approach to addressing data scarcity and as a foundation for surgical simulation, training, and robotic policy learning. The field has developed rapidly without a clear conceptual framework. This survey organizes the 2024-2026 literature into three categories: unconditional generation, conditional generation, and world modeling generation, revealing a fundamental shift in how the task is defined from synthesizing visually plausible frames to modeling the causal dynamics of surgical scenes. We examine the persistent gap between pixel-level fidelity and clinical plausibility, and identify generalization, physical realism, controllability, and interpretability as bottlenecks. We further summarize experimental results of representative methods on public datasets to provide a quantitative reference for the field. This survey provides a structured overview of the current state and open challenges, offering a reference for researchers working at the intersection of intelligent perception, multi-modal fusion, generative AI, and surgical data science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。