提出检测神经音频编码器生成的心音信号的新任务与方法
Towards Detecting Neural Audio Codec Synthesized Heart Sounds

- 融合频谱特征与自监督学习特征,提升检测能力
- 新数据集CARDIOFAKE包含真实与合成心音信号
- 适合医疗安全、生物信号伪造检测方向研究者
本文提出合成心音信号检测(SHAC)任务,旨在识别由神经音频编码器(NACs)生成的心音图(PCGs)。为推动该方向研究,我们发布了首个基准数据集CARDIOFAKE,包含真实与编码器合成的PCGs。我们对频谱表示(如MFCC、LFCC)和自监督学习表示(如WavLM)进行了基准测试。进一步提出GROOT融合框架,整合频谱与SSL特征以发挥其互补性。实验表明,结合MFCC与WavLM的GROOT达到当前最优性能,优于单一表示及主流基线。
原文摘要 · Abstract (English)
In this paper, we introduce Synthetic Heart Sound Detection (SHAC), a task aimed at identifying phonocardiograms (PCGs) synthesized using neural audio codecs (NACs). To facilitate research in this direction, we release CARDIOFAKE, the first benchmark dataset for SHAC containing both real and codec-synthesized PCGs. We benchmark spectral representations (MFCC, LFCC) and self-supervised learning (SSL) representations (e.g., WavLM) for the task. Furthermore, we propose GROOT, a fusion framework that integrates spectral and SSL features for leveraging their complementary behavior. Experiments show that GROOT, combining MFCC and WavLM, achieves state-of-the-art performance, outperforming individual representations and competitive baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。