用分层结构生成更真实的手术视频,让动作和阶段保持一致。
HieraSurg: Hierarchy-Aware Diffusion Model for Surgical Video Generation
- 分两阶段生成:先预测语义变化,再融合细节纹理。
- 在胆囊切除数据集上显著优于已有方法,支持高帧率输出。
- 适合需要精准模拟手术过程的医疗训练与研究场景。
手术视频生成已成为扩散模型在通用视频生成成功后的一个有前景的研究方向。尽管现有方法实现了高质量视频生成,但多数为无条件生成,无法保持手术动作与阶段的一致性,缺乏对手术内容的理解和细粒度引导,难以实现真实模拟。为此,我们提出HieraSurg,一种分层感知的手术视频生成框架,包含两个专用扩散模型。给定手术阶段和初始帧,HieraSurg首先通过分割预测模型预测未来的粗粒度语义变化;第二阶段模型则在时间分割图基础上叠加细粒度视觉特征,实现有效纹理渲染与语义信息融合。该方法在多个抽象层级(手术阶段、动作三元组、全景分割图)利用手术信息。在胆囊切除手术视频生成任务上的实验表明,模型在定量与定性指标上均显著优于现有工作,展现出强泛化能力,可生成更高帧率视频。当提供已有分割图时,模型表现出极佳的细粒度一致性,显示出其在实际手术应用中的潜力。
原文摘要 · Abstract (English)
Surgical Video Synthesis has emerged as a promising research direction following the success of diffusion models in general-domain video generation. Although existing approaches achieve high-quality video generation, most are unconditional and fail to maintain consistency with surgical actions and phases, lacking the surgical understanding and fine-grained guidance necessary for factual simulation. We address these challenges by proposing HieraSurg, a hierarchy-aware surgical video generation framework consisting of two specialized diffusion models. Given a surgical phase and an initial frame, HieraSurg first predicts future coarse-grained semantic changes through a segmentation prediction model. The final video is then generated by a second-stage model that augments these temporal segmentation maps with fine-grained visual features, leading to effective texture rendering and integration of semantic information in the video space. Our approach leverages surgical information at multiple levels of abstraction, including surgical phase, action triplets, and panoptic segmentation maps. The experimental results on Cholecystectomy Surgical Video Generation demonstrate that the model significantly outperforms prior work both quantitatively and qualitatively, showing strong generalization capabilities and the ability to generate higher frame-rate videos. The model exhibits particularly fine-grained adherence when provided with existing segmentation maps, suggesting its potential for practical surgical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。