用扩散模型生成带标注的导丝视频,减少人工标注需求
Label-Efficient Data Augmentation with Video Diffusion Models for Guidewire Segmentation in Cardiac Fluoroscopy
- 基于视频扩散模型,分离场景与运动分布生成新视频
- 合成视频使导丝分割精度提升,验证了生成数据有效性
- 适合医学图像标注稀缺场景,尤其适用于介入手术辅助
在介入性心脏造影视频中准确分割导丝对计算机辅助导航至关重要。尽管深度学习方法在导丝分割上表现优异,但其泛化能力依赖大量标注数据,凸显了标注成本高的问题。为此,我们提出分割引导的帧一致性视频扩散模型(SF-VD),用于生成大规模带标注的造影视频,以扩充导丝分割网络的训练数据。SF-VD通过独立建模场景分布与运动分布:先根据输入掩码生成含导丝的2D造影图像,再逐步生成后续帧,利用帧一致性策略保证帧间连贯性。此外,分割引导机制通过调整导丝对比度,实现合成图像中导丝可见性的多样性。在造影数据集上的评估表明,生成视频质量优越,且显著提升了导丝分割性能。
原文摘要 · Abstract (English)
The accurate segmentation of guidewires in interventional cardiac fluoroscopy videos is crucial for computer-aided navigation tasks. Although deep learning methods have demonstrated high accuracy and robustness in wire segmentation, they require substantial annotated datasets for generalizability, underscoring the need for extensive labeled data to enhance model performance. To address this challenge, we propose the Segmentation-guided Frame-consistency Video Diffusion Model (SF-VD) to generate large collections of labeled fluoroscopy videos, augmenting the training data for wire segmentation networks. SF-VD leverages videos with limited annotations by independently modeling scene distribution and motion distribution. It first samples the scene distribution by generating 2D fluoroscopy images with wires positioned according to a specified input mask, and then samples the motion distribution by progressively generating subsequent frames, ensuring frame-to-frame coherence through a frame-consistency strategy. A segmentation-guided mechanism further refines the process by adjusting wire contrast, ensuring a diverse range of visibility in the synthesized image. Evaluation on a fluoroscopy dataset confirms the superior quality of the generated videos and shows significant improvements in guidewire segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。