用生成模型估计心脏超声视频中的射血分数,更准确且可解释。
Generative Regression for Left Ventricular Ejection Fraction Estimation from Echocardiography Video
- 提出基于扩散模型的生成式回归方法,建模射血分数的分布而非单一数值。
- 在多个数据集上优于传统方法,噪声和异常情况下的预测更稳定。
- 适合心脏病诊断辅助,尤其对复杂病例提供可解释的预测轨迹。
从超声心动图视频中估算左心室射血分数(LVEF)是一个病态逆问题。固有的噪声、伪影和视角限制导致单个视频可能对应多个合理的生理值,而非唯一真值。现有深度学习方法通常将其视为标准回归任务,最小化均方误差(MSE),迫使模型学习条件期望,但在病理情况下后验分布常为多峰或重尾时,会产生误导性预测。本文提出从确定性回归向生成式回归的范式转变,设计多模态条件得分驱动扩散模型(MCSDR),用于建模在超声视频和患者人口统计学先验条件下,LVEF 的连续后验分布。在 EchoNet-Dynamic、EchoNet-Pediatric 与 CAMUS 数据集上的大量实验表明,MCSDR 达到当前最优性能。定性分析显示,模型生成轨迹在高噪声或显著生理变异情况下表现出明显差异,为 AI 辅助诊断提供了新的可解释性层面。
原文摘要 · Abstract (English)
Estimating Left Ventricular Ejection Fraction (LVEF) from echocardiograms constitutes an ill-posed inverse problem. Inherent noise, artifacts, and limited viewing angles introduce ambiguity, where a single video sequence may map not to a unique ground truth, but rather to a distribution of plausible physiological values. Prevailing deep learning approaches typically formulate this task as a standard regression problem that minimizes the Mean Squared Error (MSE). However, this paradigm compels the model to learn the conditional expectation, which may yield misleading predictions when the underlying posterior distribution is multimodal or heavy-tailed -- a common phenomenon in pathological scenarios. In this paper, we investigate the paradigm shift from deterministic regression toward generative regression. We propose the Multimodal Conditional Score-based Diffusion model for Regression (MCSDR), a probabilistic framework designed to model the continuous posterior distribution of LVEF conditioned on echocardiogram videos and patient demographic attribute priors. Extensive experiments conducted on the EchoNet-Dynamic, EchoNet-Pediatric, and CAMUS datasets demonstrate that MCSDR achieves state-of-the-art performance. Notably, qualitative analysis reveals that the generation trajectories of our model exhibit distinct behaviors in cases characterized by high noise or significant physiological variability, thereby offering a novel layer of interpretability for AI-aided diagnosis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。