用视频模型加速扩散模型,实现超实时医学图像逆问题求解
Sequential Posterior Sampling with Diffusion Models
- 用ViViT建模帧间动态,基于前序输出初始化反向扩散
- 推理速度提升25倍,相同质量下仅需1/25迭代次数
- 适合高帧率医学成像等需实时推理的场景
扩散模型在建模复杂分布和有效后验采样方面表现优异,但其迭代特性导致计算成本高,难以应用于超声成像等实时序列逆问题。鉴于序列帧间存在强时序结构,我们提出一种新方法,通过建模过渡动态来提升条件图像生成中序列扩散后验采样的效率。该方法利用基于先前扩散输出的视频视觉变压器(ViViT)过渡模型,将反向扩散轨迹初始化于更低噪声水平,显著减少收敛所需迭代次数。我们在真实世界高帧率心脏超声图像数据集上验证了该方法的有效性,结果表明,在保持与完整扩散轨迹相同性能的前提下,推理速度提升25倍,实现超实时后验采样。此外,过渡模型在严重运动情况下可使PSNR提升最高达8%。该方法为扩散模型在成像及其他需要实时推理领域的应用开辟了新可能。
原文摘要 · Abstract (English)
Diffusion models have quickly risen in popularity for their ability to model complex distributions and perform effective posterior sampling. Unfortunately, the iterative nature of these generative models makes them computationally expensive and unsuitable for real-time sequential inverse problems such as ultrasound imaging. Considering the strong temporal structure across sequences of frames, we propose a novel approach that models the transition dynamics to improve the efficiency of sequential diffusion posterior sampling in conditional image synthesis. Through modeling sequence data using a video vision transformer (ViViT) transition model based on previous diffusion outputs, we can initialize the reverse diffusion trajectory at a lower noise scale, greatly reducing the number of iterations required for convergence. We demonstrate the effectiveness of our approach on a real-world dataset of high frame rate cardiac ultrasound images and show that it achieves the same performance as a full diffusion trajectory while accelerating inference 25$\times$, enabling real-time posterior sampling. Furthermore, we show that the addition of a transition model improves the PSNR up to 8\% in cases with severe motion. Our method opens up new possibilities for real-time applications of diffusion models in imaging and other domains requiring real-time inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。