arXiv:2605.16530cs.CV2026-05被引 1

用符号与扩散模型结合,实现高真实感白内障手术模拟。

SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation

论文配图:SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation
图 1 · 摘自论文原文
  • 分离运动与视觉:符号系统建模器械-组织交互,扩散模型生成逼真画面。
  • 通过真实视频逆向训练,实现从仿真到真实的视频风格迁移。
  • 可泛化到未见几何结构,提升阶段识别精度,适合手术训练与智能系统开发。

真实感手术模拟对新手医生培训及自主代理开发至关重要。世界模型可通过当前观测与手术操作预测未来患者状态,从而扩展模拟环境。然而,现有先进方法难以满足临床应用的关键要求,如视觉真实感、物理合理的交互以及训练分布外场景的模拟能力。为此,我们提出SWoMo,一种用于白内障手术模拟的神经符号世界模型,将运动生成与视觉真实感解耦。符号组件(基于规则的模拟器与场景图表示)建模运动动力学与器械-组织交互,而扩散模型生成包含纹理和组织形变的真实视觉外观。我们提出逆向配对策略,利用真实手术视频重建仿真视频,获得成对数据,用于训练视频扩散模型实现从仿真到真实的反向转换。实验表明,该模型在定性与定量指标上均优于现有方法。结果显示,该模拟器满足关键标准:可泛化至未见交互几何结构,提升下游阶段检测性能,并支持无监督视频风格迁移。代码、数据与模型权重已公开于:https://ssharvienkumar.github.io/SWoMo/

原文摘要 · Abstract (English)

Realistic surgical simulation plays a crucial role in training novice surgeons and in the development of autonomous agents. World models can scale such simulation environments to realistic and diverse procedures by predicting future patient states conditioned on current observations and surgical actions. However, current state-of-the-art approaches often fail to satisfy key criteria required for clinical applicability, including visual realism, physically grounded interactions, and the ability to simulate scenarios beyond the training distribution. Hence, we introduce SWoMo, a neuro-symbolic world model for cataract surgery simulation that decouples motion generation from visual realism. The symbolic component, consisting of a rule-based simulator and scene graph representations, models motion dynamics and tool-tissue interactions, while a diffusion model produces realistic visual appearance, including textures and tissue deformations. We propose an inverse pairing strategy that reconstructs real surgical videos in the simulator to obtain paired simulated and real videos, which are then used to train our video diffusion model for the reverse objective of sim-to-real translation. Our experiments show both qualitative and quantitative improvements over prior work. We demonstrate that our simulator further satisfies the key criteria, including generalisation to unseen interaction geometries, improvements in downstream phase detection, and unsupervised video style transfer. The code, data, and model weights are available at: https://ssharvienkumar.github.io/SWoMo/

手术模拟神经符号扩散模型医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。