用轻量信号控制生成逼真手术视频,解决数据少与仿真难问题
SAW: Toward a Surgical Action World Model via Controllable and Scalable Video Generation
- 通过语言、场景、组织掩码和工具轨迹四类轻量信号控制视频生成
- 在12,044段腹腔镜视频上训练,时序一致性指标提升至199.19(原546.82)
- 适合手术人工智能与虚拟仿真领域,尤其可用于罕见动作数据增强
具备精确控制器械-组织交互能力的手术世界模型,可应对外科AI与仿真中的核心挑战——数据稀缺、罕见事件合成及仿真到现实的鸿沟。然而,现有视频生成方法需昂贵标注或复杂结构中间变量作为推理条件,限制了可扩展性;其他方法在复杂腹腔镜场景中时序一致性差且真实感不足。本文提出外科动作世界模型(SAW),通过基于四类轻量信号(语言提示、参考手术场景、组织可及性掩码、2D工具尖端轨迹)的条件视频扩散模型实现外科动作合成。该模型将视频到视频扩散重构为轨迹条件化手术动作生成,并在自建的12,044段腹腔镜视频数据集上微调,采用深度一致性损失保证几何合理性,无需推理时提供深度图。SAW在测试集上达到当前最优时序一致性(CD-FVD: 199.19 vs. 546.82)与强视觉质量。进一步实验表明:(a) 在外科AI中,使用SAW生成数据增强罕见动作,使真实测试集上的动作识别准确率从20.93%提升至43.14%(剪切动作从0.00%升至8.33%);(b) 在手术仿真中,基于模拟器导出的轨迹点生成工具-组织交互视频,构建高保真仿真引擎。
原文摘要 · Abstract (English)
A surgical world model capable of generating realistic surgical action videos with precise control over tool-tissue interactions can address fundamental challenges in surgical AI and simulation -- from data scarcity and rare event synthesis to bridging the sim-to-real gap for surgical automation. However, current video generation methods, the very core of such surgical world models, require expensive annotations or complex structured intermediates as conditioning signals at inference, limiting their scalability. Other approaches exhibit limited temporal consistency across complex laparoscopic scenes and do not possess sufficient realism. We propose Surgical Action World (SAW) -- a step toward surgical action world modeling through video diffusion conditioned on four lightweight signals: language prompts encoding tool-action context, a reference surgical scene, tissue affordance mask, and 2D tool-tip trajectories. We design a conditional video diffusion approach that reformulates video-to-video diffusion into trajectory-conditioned surgical action synthesis. The backbone diffusion model is fine-tuned on a custom-curated dataset of 12,044 laparoscopic clips with lightweight spatiotemporal conditioning signals, leveraging a depth consistency loss to enforce geometric plausibility without requiring depth at inference. SAW achieves state-of-the-art temporal consistency (CD-FVD: 199.19 vs. 546.82) and strong visual quality on held-out test data. Furthermore, we demonstrate its downstream utility for (a) surgical AI, where augmenting rare actions with SAW-generated videos improves action recognition (clipping F1-score: 20.93% to 43.14%; cutting: 0.00% to 8.33%) on real test data, and (b) surgical simulation, where rendering tool-tissue interaction videos from simulator-derived trajectory points toward a visually faithful simulation engine.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。