用单步生成技术高效合成可控的心脏超声视频,提升医疗数据模拟效率。
EchoLVFM: One-Step Video Generation via Latent Flow Matching for Echocardiogram Synthesis
- 在潜空间中通过流匹配实现单步生成,大幅提升采样速度。
- 生成视频精确控制射血分数,专家判别准确率达57.9%接近随机水平。
- 支持不完整序列重建,适合真实临床数据的异构场景使用。
心脏超声广泛用于评估心功能,其中左心室射血分数(EF)是诊断与管理的核心指标。能够显式控制此类参数并生成逼真超声视频的生成模型,在数据增强、反事实分析和专科培训中具有重要价值。然而,现有方法通常依赖计算成本高的多步采样和激进的时间归一化,限制了效率与在异构真实数据上的适用性。本文提出EchoLVFM,一种基于潜空间流匹配的一步式可控心脏超声视频生成框架。该方法在单次推理中生成时间连贯视频,相比多步流基线实现约50倍的采样效率提升,同时保持视觉保真度。模型支持全局临床变量条件控制,可精准调节EF,并能从部分观测序列中重构及生成反事实视频。采用掩码条件策略,消除固定长度约束,避免短序列被丢弃。我们在CAMUS数据集上进行单帧条件下的挑战性评估。定量与定性结果表明,生成视频质量具有竞争力,EF控制准确,专家判别准确率为57.9%,接近随机水平。结果表明,高效的单步流匹配可实现高保真、可控的心脏超声视频合成。代码已开源:https://github.com/EngEmmanuel/EchoLVFM
原文摘要 · Abstract (English)
Echocardiography is widely used for assessing cardiac function, where clinically meaningful parameters such as left-ventricular ejection fraction (EF) play a central role in diagnosis and management. Generative models capable of synthesising realistic echocardiogram videos with explicit control over such parameters are valuable for data augmentation, counterfactual analysis, and specialist training. However, existing approaches typically rely on computationally expensive multi-step sampling and aggressive temporal normalisation, limiting efficiency and applicability to heterogeneous real-world data. We introduce EchoLVFM, a one-step latent video flow-matching framework for controllable echocardiogram generation. Operating in the latent space, EchoLVFM synthesises temporally coherent videos in a single inference step, achieving a $\mathbf{\sim 50\times}$ improvement in sampling efficiency compared to multi-step flow baselines while maintaining visual fidelity. The model supports global conditioning on clinical variables, demonstrated through precise control of EF, and enables reconstruction and counterfactual generation from partially observed sequences. A masked conditioning strategy further removes fixed-length constraints, allowing shorter sequences to be retained rather than discarded. We evaluate EchoLVFM on the CAMUS dataset under challenging single-frame conditioning. Quantitative and qualitative results demonstrate competitive video quality, strong EF adherence, and 57.9% discrimination accuracy by expert clinicians which is close to chance. These findings indicate that efficient, one-step flow matching can enable practical, controllable echocardiogram video synthesis without sacrificing fidelity. Code available at: https://github.com/EngEmmanuel/EchoLVFM
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。