用临床数据生成真实心音信号,还能精准控制病理特征。
H-LDM: Hierarchical Latent Diffusion Models for Controllable and Interpretable PCG Synthesis from Clinical Metadata
- 分层潜空间分离心律、心音与杂音,实现生理解耦。
- 生成心音信号在关键指标上达9.7分(弗雷切特距离),临床有效性超87%。
- 适合心脏病诊断模型训练,尤其对罕见病提升显著。
心音图(PCG)分析对心血管疾病诊断至关重要,但标注的病理数据稀缺制约了人工智能系统的发展。为解决此问题,我们提出H-LDM——一种基于结构化临床元数据生成临床准确且可控心音信号的分层潜扩散模型。该方法包含:(1) 多尺度变分自编码器,学习生理解耦的潜空间,分离心律、心音与杂音;(2) 基于丰富临床元数据的分层文本到生物信号生成管道,实现对17种不同病理条件的细粒度控制;(3) 由新型医学注意力模块引导的可解释扩散过程。在PhysioNet CirCor数据集上的实验表明,该模型达到9.7的弗雷切特音频距离,属性解耦分数达92%,经心脏科医生验证的临床有效性为87.1%。使用合成数据增强诊断模型后,罕见病分类准确率提升11.3%。H-LDM为心脏诊断中的数据增强开辟了新路径,弥合数据稀缺与可解释临床洞察之间的鸿沟。
原文摘要 · Abstract (English)
Phonocardiogram (PCG) analysis is vital for cardiovascular disease diagnosis, yet the scarcity of labeled pathological data hinders the capability of AI systems. To bridge this, we introduce H-LDM, a Hierarchical Latent Diffusion Model for generating clinically accurate and controllable PCG signals from structured metadata. Our approach features: (1) a multi-scale VAE that learns a physiologically-disentangled latent space, separating rhythm, heart sounds, and murmurs; (2) a hierarchical text-to-biosignal pipeline that leverages rich clinical metadata for fine-grained control over 17 distinct conditions; and (3) an interpretable diffusion process guided by a novel Medical Attention module. Experiments on the PhysioNet CirCor dataset demonstrate state-of-the-art performance, achieving a Fréchet Audio Distance of 9.7, a 92% attribute disentanglement score, and 87.1% clinical validity confirmed by cardiologists. Augmenting diagnostic models with our synthetic data improves the accuracy of rare disease classification by 11.3\%. H-LDM establishes a new direction for data augmentation in cardiac diagnostics, bridging data scarcity with interpretable clinical insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。