arXiv:2510.05922eess.AS2025-10

MFCCs其实包含语音韵律信息,可提升语音模型表现

Revisiting MFCCs: Evidence for Spectral-Prosodic Coupling

  • 通过统计检验发现MFCC与音量、基频、发音特征相关
  • 三类韵律特征与MFCC的独立性在统计上不成立
  • 适合语音识别与分析模型的改进研究者参考

梅尔频率倒谱系数(MFCCs)是语音处理中的重要特征。本研究挑战了长期以来认为MFCC缺乏时间信息的观点,通过零假设显著性检验框架,系统评估了MFCC与能量、基频(F0)和发音(voicing)三类韵律特征之间的统计独立性。结果表明,MFCC与任一韵律特征独立在统计上均不成立。该发现说明MFCC本身蕴含丰富的韵律信息,有助于未来语音分析与识别模型的设计。

原文摘要 · Abstract (English)

Mel-frequency cepstral coefficients (MFCCs) are an important feature in speech processing. A deeper understanding of their properties can contribute to the work that is being done with both classical and deep learning models. This study challenges the long-held assumption that MFCCs lack relevant temporal information by investigating their relationship with speech prosody. Using a null hypothesis significance testing framework, a systematic assessment is made about the statistical independence between MFCCs and the three prosodic features: energy, fundamental frequency (F0), and voicing. The results demonstrate that it is statistically implausible that the MFCCs are independent of any of these three prosodic features. This finding suggests that MFCCs inherently carry valuable prosodic information, which can inform the design of future models in speech analysis and recognition.

语音处理特征提取韵律分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。