arXiv:2510.24519cs.SDcs.AI2025-10被引 5

提出时域梅尔小波系数,兼顾听觉感知与高效计算。

Audio Signal Processing Using Time Domain Mel-Frequency Wavelet Coefficient

  • 在时域融合梅尔尺度与小波变换,直接提取时间-频率特征。
  • 相比传统方法降低计算开销,提升音频处理效率。
  • 适合实时语音识别与资源受限设备上的音频分析。

语音信号处理中,特征提取至关重要。梅尔频率倒谱系数(MFCC)广泛应用于说话人与语音识别,因其模拟了人耳的滤波特性。但其仅提供频率信息,缺乏时间定位。小波变换具备灵活的时间-频率窗口,适合分析非平稳语音信号,但其均匀频率尺度导致低频分辨率差,且与人耳感知不一致。因此,需融合两者优势。现有方法在梅尔尺度滤波后叠加小波变换,计算复杂度高。本文提出时域梅尔频率小波系数(TMFWC)方法,直接在时域结合小波思想提取梅尔尺度特征,减少时频转换开销与小波提取复杂度。将该技术与水库计算结合,显著提升了音频信号处理效率。

原文摘要 · Abstract (English)

Extracting features from the speech is the most critical process in speech signal processing. Mel Frequency Cepstral Coefficients (MFCC) are the most widely used features in the majority of the speaker and speech recognition applications, as the filtering in this feature is similar to the filtering taking place in the human ear. But the main drawback of this feature is that it provides only the frequency information of the signal but does not provide the information about at what time which frequency is present. The wavelet transform, with its flexible time-frequency window, provides time and frequency information of the signal and is an appropriate tool for the analysis of non-stationary signals like speech. On the other hand, because of its uniform frequency scaling, a typical wavelet transform may be less effective in analysing speech signals, have poorer frequency resolution in low frequencies, and be less in line with human auditory perception. Hence, it is necessary to develop a feature that incorporates the merits of both MFCC and wavelet transform. A great deal of studies are trying to combine both these features. The present Wavelet Transform based Mel-scaled feature extraction methods require more computation when a wavelet transform is applied on top of Mel-scale filtering, since it adds extra processing steps. Here we are proposing a method to extract Mel scale features in time domain combining the concept of wavelet transform, thus reducing the computational burden of time-frequency conversion and the complexity of wavelet extraction. Combining our proposed Time domain Mel frequency Wavelet Coefficient(TMFWC) technique with the reservoir computing methodology has significantly improved the efficiency of audio signal processing.

语音处理特征提取小波变换高效计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。