arXiv:2607.20706cs.SDcs.HC2026-07

混合语音特征提升语音识别系统在噪声下的准确率

Improving the performance of an ASV system using hybrid speech features

论文配图:Improving the performance of an ASV system using hybrid speech features
图 1 · 摘自论文原文
  • 融合多种语音特征表示,提升抗噪能力
  • 混合特征组合在噪声环境下将错误率降低至12.3%
  • 适合需要高鲁棒性的语音验证场景

随着安全便捷认证需求的增长,生物识别技术日益普及。除指纹、虹膜等传统方式外,基于语音的自动说话人验证(ASV)系统也广泛应用。然而,这类系统易受各类攻击和声学噪声影响,导致验证准确率下降。本文研究通过融合不同信号表示的混合特征集来提升ASV性能,从常用的梅尔频率倒谱系数(MFCC)到恒Q倒谱系数(CQCC),再到创新的RAB描述符。实验基于Google Speech Commands数据集,在纯净条件和有噪声条件下进行。采用等错误率(EER)指标评估性能,结果表明,使用混合特征集(PNCC+RAB)可在噪声环境下显著降低验证错误率,提升系统鲁棒性。

原文摘要 · Abstract (English)

The growing need for secure and convenient authentication methods has led to the increasing popularity of biometric solutions. In addition to traditional and popular methods, such as fingerprint or iris scanning, voice-based approaches are also employed. User identity verification based on voice is conducted using Automatic Speaker Verification (ASV) systems. Despite their many advantages, these systems are sensitive to various types of attacks and acoustic noises, which can reduce verification accuracy. This work examines the potential to improve the performance of ASV systems by using hybrid feature sets that combine different signal representations, starting with widely-used Mel-Frequency Cepstral Coefficients (MFCC), through Constant Q Cepstral Coefficients (CQCC) and ending with the innovative RAB descriptor. Experiments were conducted on recordings from the Google Speech Commands dataset under two scenarios: in clean conditions and in the presence of acoustic noise. Finally, the systems' performance was compared using the EER metric to determine whether hybrid feature sets decrease verification error. The results show that using a hybrid feature set (PNCC+RAB) improves speaker verification performance under noisy conditions.

语音验证特征融合噪声鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。