用扩散模型学习微创手术缝合动作,实现高保真动态模拟。
Towards Suturing World Models: Learning Predictive Models for Robotic Surgical Tasks
- 基于标注手术视频,训练扩散模型捕捉缝合动作的时空动态。
- 生成分辨率≥768x512、时长≥49帧的高保真动作序列,可区分理想与非理想操作。
- 适用于手术训练模拟器、技能评估系统及自主手术系统研究。
我们提出一种基于扩散的生成模型,通过监督学习标注的腹腔镜手术视频,捕捉精细缝合动作的时空动态。构建了一个包含约2000段剪辑的数据集,将手术动作细分为针头定位、对准、穿刺和退出等子动作类别,涵盖理想与非理想执行方式。我们微调了两种先进视频扩散模型(LTX-Video 和 HunyuanVideo),以生成分辨率≥768x512、帧数≥49的高质量手术动作序列。采用低秩适配(LoRA)与全模型微调两种方法进行训练。实验表明,这些世界模型能有效捕捉缝合动态,具备区分理想与非理想操作的能力,为改进手术训练模拟器、技能评估工具及自主手术系统提供基础。模型已开源,供测试与后续研究使用。
原文摘要 · Abstract (English)
We introduce specialized diffusion-based generative models that capture the spatiotemporal dynamics of fine-grained robotic surgical sub-stitch actions through supervised learning on annotated laparoscopic surgery footage. The proposed models form a foundation for data-driven world models capable of simulating the biomechanical interactions and procedural dynamics of surgical suturing with high temporal fidelity. Annotating a dataset of $\sim2K$ clips extracted from simulation videos, we categorize surgical actions into fine-grained sub-stitch classes including ideal and non-ideal executions of needle positioning, targeting, driving, and withdrawal. We fine-tune two state-of-the-art video diffusion models, LTX-Video and HunyuanVideo, to generate high-fidelity surgical action sequences at $\ge$768x512 resolution and $\ge$49 frames. For training our models, we explore both Low-Rank Adaptation (LoRA) and full-model fine-tuning approaches. Our experimental results demonstrate that these world models can effectively capture the dynamics of suturing, potentially enabling improved training simulators, surgical skill assessment tools, and autonomous surgical systems. The models also display the capability to differentiate between ideal and non-ideal technique execution, providing a foundation for building surgical training and evaluation systems. We release our models for testing and as a foundation for future research. Project Page: https://mkturkcan.github.io/suturingmodels/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。