用文本生成高保真心脏影像,实现时序一致与病理多样。
Temporally Consistent and Controllable Video Generation of 2D Cine CMR via Latent Space Motion Modeling

- 分离解剖结构与运动轨迹,在潜在空间建模心脏动态
- 生成序列FID达31.68,文本对齐CLIP得分31.04
- 适合医学数据增强与个性化心功能模拟
心脏电影磁共振(Cine CMR)是评估心脏功能的金标准,但公开数据集稀缺限制了数据驱动模型的发展。为此,我们提出一种生成方法,用于合成时序一致且解剖结构准确的心脏序列。所提文本到视频框架将心脏空间结构与时间运动分离:首先通过微调的扩散模型根据临床文本提示生成初始帧,控制解剖特征;随后基于心脏相位嵌入的潜在流模型生成完整心脏运动,确保空间一致性与时序可控性。模型生成的序列在解剖与病理上具有多样性,具备高时序连贯性与强提示契合度,图像真实感FID为31.68,文本-图像对齐CLIP得分为31.04。实验结果表明该方法可生成高保真、按需定制的医学数据,为数据稀缺问题提供可扩展的解决方案。
原文摘要 · Abstract (English)
Cine cardiac magnetic resonance is the gold standard for assessing cardiac function, but the scarcity of public datasets limits the development of advanced data-driven models. To address this limitation, we propose a generative method for synthesizing temporally coherent and anatomically consistent cardiac sequences. Our text-to-video framework decouples cardiac spatial structure from temporal motion. First, a fine-tuned diffusion model synthesizes an initial frame from a clinical text prompt, controlling anatomical features. Then, a latent flow model conditioned on a cardiac phase embedding generates the complete cardiac motion, ensuring spatial consistency and temporal control. Our model generates anatomically and pathologically diverse sequences with high temporal coherence and strong fidelity to input prompts, achieving a FID of 31.68 for image realism and a CLIP score of 31.04 for text-image alignment. These experimental results highlight its potential to produce high-fidelity, on-demand medical data, offering a scalable solution to data scarcity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。