用场景图控制生成高保真手术模拟图像
SurGrID: Controllable Surgical Simulation via Scene Graph to Image Diffusion
- 通过场景图编码解剖结构的空间语义关系
- 生成图像与图结构一致性提升,视觉保真度更高
- 临床专家评估验证了模拟真实性和交互可控性
手术模拟为传统外科训练提供了有力补充。然而现有工具缺乏逼真视觉效果,且依赖硬编码行为。去噪扩散模型在高质量图像生成方面表现优异,但当前最先进的条件生成方法在场景控制精度和交互性上仍有不足。本文提出SurGrID,一种基于场景图到图像扩散的手术场景生成模型。场景图编码手术场景中各组件的空间与语义信息,经由我们提出的预训练步骤,显式捕捉局部与全局上下文。该方法显著提升了生成图像的保真度及其与输入图的一致性。进一步的用户评估研究(包含临床专家)证实了模拟的真实感与可控性。结果表明,场景图可有效用于精确且交互式的扩散模型条件控制,实现高保真、可调控的手术场景生成。
原文摘要 · Abstract (English)
Surgical simulation offers a promising addition to conventional surgical training. However, available simulation tools lack photorealism and rely on hardcoded behaviour. Denoising Diffusion Models are a promising alternative for high-fidelity image synthesis, but existing state-of-the-art conditioning methods fall short in providing precise control or interactivity over the generated scenes. We introduce SurGrID, a Scene Graph to Image Diffusion Model, allowing for controllable surgical scene synthesis by leveraging Scene Graphs. These graphs encode a surgical scene's components' spatial and semantic information, which are then translated into an intermediate representation using our novel pre-training step that explicitly captures local and global information. Our proposed method improves the fidelity of generated images and their coherence with the graph input over the state-of-the-art. Further, we demonstrate the simulation's realism and controllability in a user assessment study involving clinical experts. Scene Graphs can be effectively used for precise and interactive conditioning of Denoising Diffusion Models for simulating surgical scenes, enabling high fidelity and interactive control over the generated content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。