用Mamba+光流对齐提升医学视频生成质量与效率
Optical Flow Representation Alignment Mamba Diffusion Model for Medical Video Generation
- 融合注意力机制与Mamba结构,降低计算开销同时保持高质量生成
- 通过光流表示对齐增强帧间像素关注,改善时序一致性
- 采用带频域补偿的VAE,减少医学特征在隐空间中的信息损失
医学视频生成模型有望在医疗教育、手术规划和模拟等领域产生深远影响。现有视频扩散模型多基于图像扩散架构并引入时间操作(如3D卷积和时间注意力),虽有效但过度简化限制了时空表现力且计算开销大。为此,我们提出医学仿真视频生成器MedSora,包含三个关键组件:(i) 结合注意力与Mamba优势的视频扩散框架,兼顾低计算负载与高质量生成;(ii) 光流表示对齐方法,隐式增强对帧间像素的关注;(iii) 带频域补偿的视频变分自编码器(VAE),缓解像素空间到隐空间转换过程中的医学特征信息丢失。大量实验与应用表明,MedSora在生成医学视频方面视觉质量优于最先进基线方法。更多结果与代码见https://wongzbb.github.io/MedSora
原文摘要 · Abstract (English)
Medical video generation models are expected to have a profound impact on the healthcare industry, including but not limited to medical education and training, surgical planning, and simulation. Current video diffusion models typically build on image diffusion architecture by incorporating temporal operations (such as 3D convolution and temporal attention). Although this approach is effective, its oversimplification limits spatio-temporal performance and consumes substantial computational resources. To counter this, we propose Medical Simulation Video Generator (MedSora), which incorporates three key elements: i) a video diffusion framework integrates the advantages of attention and Mamba, balancing low computational load with high-quality video generation, ii) an optical flow representation alignment method that implicitly enhances attention to inter-frame pixels, and iii) a video variational autoencoder (VAE) with frequency compensation addresses the information loss of medical features that occurs when transforming pixel space into latent features and then back to pixel frames. Extensive experiments and applications demonstrate that MedSora exhibits superior visual quality in generating medical videos, outperforming the most advanced baseline methods. Further results and code are available at https://wongzbb.github.io/MedSora
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。