arXiv:2508.17914cs.CL2025-08

分析Wav2Vec中声学特征对元音的表征能力,揭示深层网络更优的语音表示。

Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs

  • 用MFCC和带共振峰的MFCC对比CNN提取的特征
  • 在TIMIT语料上测试前端-后端元音分类准确率
  • 发现高层卷积层特征表现更优,适合语音表征研究

自动语音识别借助自监督学习实现了从原始音频直接提取特征。在Wav2Vec中,首先通过卷积神经网络(CNN)将音频转换为特征向量,再由变压器(Transformer)处理。本研究基于TIMIT语料库,考察了单母音(monophthong vowels)在该模型中卷积层提取的信息。通过训练支持向量机(SVM)分类器进行前后元音判别任务,比较了传统梅尔频率倒谱系数(MFCCs)、含共振峰信息的MFCCs以及CNN激活值的表现,并以分类准确率评估其语音表征能力。

原文摘要 · Abstract (English)

Automatic Speech Recognition has advanced with self-supervised learning, enabling feature extraction directly from raw audio. In Wav2Vec, a CNN first transforms audio into feature vectors before the transformer processes them. This study examines CNN-extracted information for monophthong vowels using the TIMIT corpus. We compare MFCCs, MFCCs with formants, and CNN activations by training SVM classifiers for front-back vowel identification, assessing their classification accuracy to evaluate phonetic representation.

语音表征特征提取深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。