arXiv:2509.02259cs.SDcs.LG2025-09中稿 · WOCCI2025

用预训练语音模型分析婴儿哭声,仅靠少量标注数据就可识别哭声特征。

Speech transformer models for extracting information from baby cries

  • 用五个预训练语音模型提取婴儿哭声的潜在表征。
  • 在8个数据集上实现超过90%的分类准确率,成功识别哭声中的身份与发声不稳信息。
  • 为情绪识别等类似任务提供模型设计参考,适合小样本场景研究者。

利用预训练语音模型的潜在表征进行迁移学习,在标注数据稀缺的任务中表现优异。然而,这些模型在非语音数据上的适用性及其表征中编码的特定声学特性仍不明确。本研究评估了五个预训练语音模型在八个婴儿哭声数据集上的表现,涵盖960名婴儿的115小时音频。针对每个数据集,我们测试了各模型在所有可用分类任务中的潜在表征效果。结果表明,这些模型的潜在表征能有效分类人类婴儿哭声,并编码与发声源不稳定性及哭婴身份相关的关键信息。此外,对模型架构和训练策略的对比分析,为未来面向类似任务(如情绪检测)的模型设计提供了重要启示。

原文摘要 · Abstract (English)

Transfer learning using latent representations from pre-trained speech models achieves outstanding performance in tasks where labeled data is scarce. However, their applicability to non-speech data and the specific acoustic properties encoded in these representations remain largely unexplored. In this study, we investigate both aspects. We evaluate five pre-trained speech models on eight baby cries datasets, encompassing 115 hours of audio from 960 babies. For each dataset, we assess the latent representations of each model across all available classification tasks. Our results demonstrate that the latent representations of these models can effectively classify human baby cries and encode key information related to vocal source instability and identity of the crying baby. In addition, a comparison of the architectures and training strategies of these models offers valuable insights for the design of future models tailored to similar tasks, such as emotion detection.

语音模型婴儿哭声迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。