arXiv:2510.09344cs.SDeess.AS2025-10中稿 · Interspeech 2026被引 2

构建真实场景下中文老年人语音数据集,支持语音识别与说话人分析研究

WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations

  • 从网络视频采集真实环境语音,人工精细标注转写与声学特征
  • 包含120小时语音、68位老人、9类方言口音,年龄跨度55-93岁
  • 为老年语音识别与声纹分析提供高现实性基准数据集

由于年龄相关的语言变化(如语速减慢、嗓音震颤),老年人语音在自动处理中面临独特挑战。现有中文语音数据集多在受控环境下录制,缺乏多样性与真实场景适用性。为填补这一空白,我们提出WildElder——一个从在线视频中收集的普通话老年人语音语料库,并配备精细的人工标注,包括转写文本、说话人年龄、性别和口音强度。结合真实场景数据与专家标注,WildElder可支撑鲁棒的语音识别与说话人画像研究。实验结果揭示了老年语音识别的困难性,同时验证了WildElder作为新基准的挑战性潜力。数据集与代码已开源:https://github.com/NKU-HLT/WildElder。

原文摘要 · Abstract (English)

Elderly speech poses unique challenges for automatic processing due to age-related changes such as slower articulation and vocal tremors. Existing Chinese datasets are mostly recorded in controlled environments, limiting their diversity and real-world applicability. To address this gap, we present WildElder, a Mandarin elderly speech corpus collected from online videos and enriched with fine-grained manual annotations, including transcription, speaker age, gender, and accent strength. Combining the realism of in-the-wild data with expert curation, WildElder enables robust research on automatic speech recognition and speaker profiling. Experimental results reveal both the difficulties of elderly speech recognition and the potential of WildElder as a challenging new benchmark. The dataset and code are available at https://github.com/NKU-HLT/WildElder.

语音识别老年语音真实数据标注数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。