SurgSora可精准生成手术视频,支持用户控制器械运动轨迹。
SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
- 通过自预测物体特征与深度信息,提升视频生成的细节真实度。
- 在单帧输入下实现高保真、可控制的手术视频生成,优于现有方法。
- 适合医学教育、手术模拟场景,专家评价认为视频极具真实感。
手术视频生成有助于医学教育与研究,但现有方法缺乏精细运动控制和真实感。本文提出SurgSora,一种从单张输入图像和用户指定运动指令生成高保真、可控制手术视频的框架。不同于以往将物体一视同仁或依赖真实分割掩码的方法,SurgSora利用自预测物体特征与深度信息,优化RGB外观与光流以实现精确合成。其包含三个核心模块:(1) 双语义注入器,提取对象特定的RGB-D特征与分割线索以增强空间表征;(2) 解耦光流映射器,融合多尺度光流与语义特征,实现逼真的运动动态;(3) 轨迹控制器,估计稀疏光流并支持用户引导物体移动。通过在Stable Video Diffusion中引入这些增强特征,SurgSora在视觉真实性与可控性方面达到当前最佳水平,经大量定量与定性对比验证。与外科专家合作的人类评估进一步证实生成视频高度真实,展现出在手术训练与教育中的巨大潜力。项目地址:https://surgsora.github.io/surgsora.github.io。
原文摘要 · Abstract (English)
Surgical video generation can enhance medical education and research, but existing methods lack fine-grained motion control and realism. We introduce SurgSora, a framework that generates high-fidelity, motion-controllable surgical videos from a single input frame and user-specified motion cues. Unlike prior approaches that treat objects indiscriminately or rely on ground-truth segmentation masks, SurgSora leverages self-predicted object features and depth information to refine RGB appearance and optical flow for precise video synthesis. It consists of three key modules: (1) the Dual Semantic Injector, which extracts object-specific RGB-D features and segmentation cues to enhance spatial representations; (2) the Decoupled Flow Mapper, which fuses multi-scale optical flow with semantic features for realistic motion dynamics; and (3) the Trajectory Controller, which estimates sparse optical flow and enables user-guided object movement. By conditioning these enriched features within the Stable Video Diffusion, SurgSora achieves state-of-the-art visual authenticity and controllability in advancing surgical video synthesis, as demonstrated by extensive quantitative and qualitative comparisons. Our human evaluation in collaboration with expert surgeons further demonstrates the high realism of SurgSora-generated videos, highlighting the potential of our method for surgical training and education. Our project is available at https://surgsora.github.io/surgsora.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。