用场景图控制手术视频生成,实现精细操作与真实感兼具。
SG2VID: Scene Graphs Enable Fine-Grained Control for Video Synthesis
- 基于场景图的扩散模型,实现工具与解剖结构的精准控制。
- 在白内障和胆囊切除术数据集上超越现有方法,支持动态布局调整。
- 适合医学仿真、数据增强及罕见手术事件生成场景。
外科手术模拟在新手外科医生培训中至关重要,可加速学习曲线并减少术中错误。然而,传统模拟工具难以提供所需的逼真视觉效果和人体解剖结构的多样性。为此,当前方法正转向基于生成模型的模拟器。但这些方法多聚焦于复杂条件以实现精确合成,忽视了精细的人类控制能力。为填补这一空白,我们提出SG2VID,首个基于扩散模型且利用场景图实现精准视频合成与细粒度人类控制的模型。我们在包含白内障和胆囊切除术的三个公开数据集上验证其性能。实验表明,SG2VID在定性和定量上均优于先前方法,能准确控制器械与解剖结构的尺寸、运动、新器械引入及整体场景布局。我们定性展示了其在生成式数据增强中的应用,并通过实验证明,使用合成视频扩充训练集可提升下游阶段检测任务的性能。最后,为展示其保持人类控制的能力,我们通过交互修改场景图,生成包含重大但罕见术中异常的新视频样本。
原文摘要 · Abstract (English)
Surgical simulation plays a pivotal role in training novice surgeons, accelerating their learning curve and reducing intra-operative errors. However, conventional simulation tools fall short in providing the necessary photorealism and the variability of human anatomy. In response, current methods are shifting towards generative model-based simulators. Yet, these approaches primarily focus on using increasingly complex conditioning for precise synthesis while neglecting the fine-grained human control aspect. To address this gap, we introduce SG2VID, the first diffusion-based video model that leverages Scene Graphs for both precise video synthesis and fine-grained human control. We demonstrate SG2VID's capabilities across three public datasets featuring cataract and cholecystectomy surgery. While SG2VID outperforms previous methods both qualitatively and quantitatively, it also enables precise synthesis, providing accurate control over tool and anatomy's size and movement, entrance of new tools, as well as the overall scene layout. We qualitatively motivate how SG2VID can be used for generative augmentation and present an experiment demonstrating its ability to improve a downstream phase detection task when the training set is extended with our synthetic videos. Finally, to showcase SG2VID's ability to retain human control, we interact with the Scene Graphs to generate new video samples depicting major yet rare intra-operative irregularities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。