arXiv:2508.08924eess.AScs.AI2025-08被引 1

用神经编码器提升电声门图信号重建与基频提取精度

EGGCodec: A Robust Neural Encodec Framework for EGG Reconstruction and F0 Extraction

  • 通过多尺度频域损失+时域相关性损失优化信号重建
  • 基频提取误差降低至13.69赫兹,发声判断错误率下降38.2%
  • 无需生成对抗网络,训练更简单,适合语音分析研究者

本文提出EGGCodec,一种专为电声门图(EGG)信号重建与基频(F0)提取设计的鲁棒神经编码框架。通过引入多尺度频域损失函数捕捉原始与重建信号间的细微关系,并结合时域相关性损失提升泛化能力与准确性。不同于传统Encodec模型直接从特征中提取F0,EGGCodec利用重构后的EGG信号进行F0估计,更贴近真实生理对应关系。通过移除传统GAN判别器,简化训练流程且仅带来可忽略的性能损失。在常用含EGG数据集上训练后,大量实验表明,EGGCodec显著优于现有F0提取方法:平均绝对误差(MAE)由14.14赫兹降至13.69赫兹,发声判断错误率(VDE)降低38.2%。消融实验验证了各模块的有效性。

原文摘要 · Abstract (English)

This letter introduces EGGCodec, a robust neural Encodec framework engineered for electroglottography (EGG) signal reconstruction and F0 extraction. We propose a multi-scale frequency-domain loss function to capture the nuanced relationship between original and reconstructed EGG signals, complemented by a time-domain correlation loss to improve generalization and accuracy. Unlike conventional Encodec models that extract F0 directly from features, EGGCodec leverages reconstructed EGG signals, which more closely correspond to F0. By removing the conventional GAN discriminator, we streamline EGGCodec's training process without compromising efficiency, incurring only negligible performance degradation. Trained on a widely used EGG-inclusive dataset, extensive evaluations demonstrate that EGGCodec outperforms state-of-the-art F0 extraction schemes, reducing mean absolute error (MAE) from 14.14 Hz to 13.69 Hz, and improving voicing decision error (VDE) by 38.2\%. Moreover, extensive ablation experiments validate the contribution of each component of EGGCodec.

信号重建基频提取语音分析神经编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。