用动态谱特征提升越南语语音识别,更抗性别差异。
Exploring Dynamic Parameters for Vietnamese Gender-Independent ASR
- 用极坐标描述频带重心变化,捕捉语音动态特性。
- 结合MFCC使越南语识别错误率显著下降。
- 新特征对男女音色差异不敏感,适合跨性别应用。
语音信号的动态特性蕴含时间信息,在自动语音识别中至关重要。本文通过极坐标参数表征频带重心频率(SSCF)在比值平面上的声学过渡,以捕捉语音动态特征并降低频谱变化影响。这些动态参数与梅尔频率倒谱系数(MFCCs)结合用于越南语语音识别,以获取更精细的频谱信息。将SSCF0作为基频(F0)的伪特征,稳健描述声调信息。实验表明,所提参数显著降低词错误率,并表现出比基线MFCC更强的性别独立性。
原文摘要 · Abstract (English)
The dynamic characteristics of speech signal provides temporal information and play an important role in enhancing Automatic Speech Recognition (ASR). In this work, we characterized the acoustic transitions in a ratio plane of Spectral Subband Centroid Frequencies (SSCFs) using polar parameters to capture the dynamic characteristics of the speech and minimize spectral variation. These dynamic parameters were combined with Mel-Frequency Cepstral Coefficients (MFCCs) in Vietnamese ASR to capture more detailed spectral information. The SSCF0 was used as a pseudo-feature for the fundamental frequency (F0) to describe the tonal information robustly. The findings showed that the proposed parameters significantly reduce word error rates and exhibit greater gender independence than the baseline MFCCs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。