分析自监督模型如何编码儿童语音中的年龄与性别信息。
How Well Do Self-Supervised Speech Models Encode Age and Gender in Children's Speech? A Layer-Wise Analysis Across Multiple Architectures
- 分层分析四个主流自监督模型的语音表征能力。
- 早期到中期层对年龄和性别信息编码最强,1-3秒语音即可准确分类。
- 结果在跨数据集、跨验证下稳定,适合儿童语音识别研究者。
自监督学习(SSL)模型已成为现代语音处理系统的核心,能够无需标注数据即可学习丰富的声学表示。尽管在成人语音上表现优异,但这些模型在儿童语音中是否有效捕捉年龄和性别等说话人属性仍不明确。由于儿童语音受生理与认知发育影响,具有更高音高、更强发音可变性及随年龄变化的声学特征,因此更具挑战性。本文对四种广泛使用的SSL模型——Wav2Vec2、HuBERT、Data2Vec和WavLM——在儿童语音中的层间表征进行了全面分析。通过在两个基准儿童语音语料库PFSTAR和CMU Kids上提取分层特征,并使用轻量级CNN进行评估,结合主成分分析(PCA)研究特征紧凑性与冗余性。实验表明,年龄与性别信息在各层分布不均,早期至中期层编码最强的副语言线索。其中,HuBERT在年龄分类上表现最佳,而Wav2Vec2和HuBERT分别在PFSTAR和CMU Kids上主导性别分类。此外,在按说话人交叉验证、层聚合和跨数据库评估下,结果依然稳健,表明对数据不平衡与领域差异具有鲁棒性。最终证明,即使仅用1–3秒短语音片段,也能实现可靠的年龄与性别分类。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) models have become a central component of modern speech processing systems, as they enable the learning of rich acoustic representations without reliance on labeled data. Despite their success on adult speech, it remains unclear how effectively these models capture speaker-related attributes such as age and gender in children's speech, which differs substantially from adult speech due to ongoing physiological and cognitive development. Higher pitch, increased articulatory variability, and age-dependent acoustic changes make children's speech a particularly challenging domain. In this work, we present a comprehensive analysis of how age and gender information is encoded across layers of four widely used SSL models: Wav2Vec2, HuBERT, Data2Vec, and WavLM. Layer-wise features are extracted and evaluated using a lightweight CNN on two benchmark children's speech corpora, PFSTAR and CMU Kids. To analyze feature compactness and redundancy, PCA is applied to identify redundancy and highlight the dimensions that contribute most to classification performance. Experimental results show that age- and gender-related information is unevenly distributed across SSL layers, with early to mid-level layers encoding the strongest paralinguistic cues. HuBERT achieves the best overall performance for age classification, while Wav2Vec2 and HuBERT lead gender classification on PFSTAR and CMU Kids, respectively. Beyond single-split evaluation, we further demonstrate that these findings remain stable under speaker-wise cross-validation, layer aggregation, and cross-database evaluation, indicating robustness to data imbalance and domain mismatch. Finally, we show that reliable age and gender classification is achievable even from short speech segments of 1--3 seconds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。