arXiv:2502.12007cs.CLcs.AI2025-02中稿 · The Conference on …被引 6

用WavLM提取语音特征,精准预测年龄、性别等人口属性。

Demographic Attributes Prediction from Speech Using WavLM Embeddings

  • 基于WavLM预训练模型提取语音嵌入,构建通用分类器。
  • 年龄预测平均误差4.94,性别识别准确率超99.81%。
  • 在多数据集上表现更优,适合语音个性化与无障碍应用。

本文提出一种基于WavLM特征的通用分类器,用于从语音中推断年龄、性别、母语、教育程度及国家等人口属性。该任务在语言学习、无障碍技术与数字取证等领域具有重要意义,有助于实现更个性化和包容性的技术应用。通过利用预训练模型提取语音嵌入,该框架识别出与人口属性相关的声学与语言特征,在多个数据集上实现年龄预测的均方绝对误差(MAE)为4.94,性别分类准确率超过99.81%。相比现有模型,系统在各任务上相对提升达30%(MAE)和10%(准确率与F1分数),得益于多样数据集与大模型的联合使用,显著增强了鲁棒性与泛化能力。本研究为语音驱动的人口统计分析提供了新视角,并奠定了未来研究基础。

原文摘要 · Abstract (English)

This paper introduces a general classifier based on WavLM features, to infer demographic characteristics, such as age, gender, native language, education, and country, from speech. Demographic feature prediction plays a crucial role in applications like language learning, accessibility, and digital forensics, enabling more personalized and inclusive technologies. Leveraging pretrained models for embedding extraction, the proposed framework identifies key acoustic and linguistic fea-tures associated with demographic attributes, achieving a Mean Absolute Error (MAE) of 4.94 for age prediction and over 99.81% accuracy for gender classification across various datasets. Our system improves upon existing models by up to relative 30% in MAE and up to relative 10% in accuracy and F1 scores across tasks, leveraging a diverse range of datasets and large pretrained models to ensure robustness and generalizability. This study offers new insights into speaker diversity and provides a strong foundation for future research in speech-based demographic profiling.

语音分析人口属性WavLM特征提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。