arXiv:2608.06770cs.AIcs.CV2026-08

提出统一多模态控制的手术世界模型,生成更真实、连贯的手术视频。

Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts

论文配图:Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts
图 1 · 摘自论文原文
  • 用分层锚点保持解剖结构一致,避免形变和器械漂移。
  • 多模态专家融合边缘、深度、光流信息,提升运动一致性。
  • 适合需要高保真手术模拟与可控生成的研究者使用。

可控手术世界模型可为手术人工智能与仿真提供生成基础,实现逼真的器械-组织交互。然而现有方法缺乏统一的多模态控制范式,直接融合异构视觉条件常导致解剖畸变、器械外观漂移及时间不一致的交互。本文提出Surg-UniWorld,一种具有多模态控制专家的统一手术世界模型。首先,从首帧外观与分层语义掩码构建分层手术锚点,以保持场景身份、解剖结构与交互边界持久性。随后,锚点相对模态专家基于共享锚点解析边缘、深度与光流证据,捕获互补的边界、几何与运动信息。多模态控制专家进一步进行贡献保留的分阶段组合,生成控制提示供Wan2.2视频扩散主干使用。为支持多模态手术世界建模,我们构建了Cholec80-SurgWAM基准数据集。大量实验表明,Surg-UniWorld在生成质量、时间一致性和多模态可控性上均显著优于现有可控视频生成方法与手术世界模型基线。

原文摘要 · Abstract (English)

Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic instrument--tissue interactions. However, existing methods lack a unified multimodal control paradigm, while direct fusion of heterogeneous visual conditions often causes anatomical distortion, instrument appearance drift, and temporally inconsistent interactions. In this work, we propose {Surg-UniWorld}, a unified surgical world model with multimodal control experts. Surg-UniWorld first constructs a {Hierarchical Surgical Anchor} from first-frame appearance and hierarchical semantic masks to preserve persistent scene identity, anatomical organization, and interaction boundaries. {Anchor-Relative Modality Experts} then interpret edge, depth, and optical-flow evidence relative to the shared anchor, capturing complementary boundary, geometric, and motion information. A {Multimodal Control Expert} further performs contribution-preserving stage-wise composition of the activated modality increments and generates control hints for the Wan2.2 video diffusion backbone. To support multimodal surgical world modeling, we further construct Cholec80-SurgWAM, a benchmark for controllable surgical video generation. Extensive experiments demonstrate that Surg-UniWorld consistently outperforms existing controllable video generation methods and surgical world-model baselines in generation quality, temporal consistency, and multimodal controllability.

手术模拟多模态控制视频生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。