分析自监督模型如何在儿童语音中编码年龄性别特征
Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech
- 分层分析四种Wav2Vec2模型,发现浅层更擅长捕捉说话人特征
- 在CMU Kids数据集上年龄分类准确率达97.14%,性别达98.20%
- 适合开发针对儿童语音的智能语音交互系统
儿童语音因音高、发音和发育差异大,在年龄与性别分类上存在挑战。尽管自监督学习(SSL)在成人语音任务中表现优异,但其对儿童语音中说话人特征的建模仍不明确。本文对四种Wav2Vec2变体在PFSTAR和CMU Kids数据集上进行了细致的层间分析。结果表明,早期层(1-7)比深层更有效捕捉说话人特异性线索,深层则逐渐聚焦于语言信息。结合主成分分析(PCA)进一步降低冗余,突出关键特征。Wav2Vec2-large-lv60模型在CMU Kids上达到97.14%(年龄)和98.20%(性别)的准确率;base-100h与large-lv60模型在PFSTAR上分别达到86.05%和95.00%。研究揭示了说话人特征在SSL模型深度中的分布规律,为构建更精准的儿童感知语音接口提供了指导。
原文摘要 · Abstract (English)
Children's speech presents challenges for age and gender classification due to high variability in pitch, articulation, and developmental traits. While self-supervised learning (SSL) models perform well on adult speech tasks, their ability to encode speaker traits in children remains underexplored. This paper presents a detailed layer-wise analysis of four Wav2Vec2 variants using the PFSTAR and CMU Kids datasets. Results show that early layers (1-7) capture speaker-specific cues more effectively than deeper layers, which increasingly focus on linguistic information. Applying PCA further improves classification, reducing redundancy and highlighting the most informative components. The Wav2Vec2-large-lv60 model achieves 97.14% (age) and 98.20% (gender) on CMU Kids; base-100h and large-lv60 models reach 86.05% and 95.00% on PFSTAR. These results reveal how speaker traits are structured across SSL model depth and support more targeted, adaptive strategies for child-aware speech interfaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。