arXiv:2503.16578cs.CLcs.SD2025-03NeurIPS被引 16

构建了面向超高龄老人的中文对话语音数据集,解决老龄化语音技术训练数据不足问题。

SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors

  • 采集101人202位参与者的55.53小时自然对话,覆盖性别、地域、年龄多样性
  • 标注多维度语音特征,支持声纹验证、语音识别等6类任务实验
  • 填补75岁以上超龄人群语音数据空白,适合老龄化语音系统研究者使用

尽管语音技术日益服务于老年人群,但现有系统因缺乏能捕捉老年人特有语音特征(如老年性发音障碍和方言差异)的训练数据,导致性能显著下降。现有老年人语音数据集中对75岁以上超高龄个体的数据稀缺,且录音方式过于简单、标注维度单一,进一步加剧了这一问题。为解决该关键数据短缺,我们提出SeniorTalk,一个精心标注的中文口语对话数据集。该数据集包含101位参与者在202人参与的101场自然对话中产生的55.53小时语音,确保性别、地域与年龄分布的策略性平衡。通过多维度详细标注,可支持广泛语音任务。我们在声纹验证、说话人分离、语音识别及语音编辑等任务上开展全面实验,为面向该年龄段的语音技术发展提供关键洞见。

原文摘要 · Abstract (English)

While voice technologies increasingly serve aging populations, current systems exhibit significant performance gaps due to inadequate training data capturing elderly-specific vocal characteristics like presbyphonia and dialectal variations. The limited data available on super-aged individuals in existing elderly speech datasets, coupled with overly simple recording styles and annotation dimensions, exacerbates this issue. To address the critical scarcity of speech data from individuals aged 75 and above, we introduce SeniorTalk, a carefully annotated Chinese spoken dialogue dataset. This dataset contains 55.53 hours of speech from 101 natural conversations involving 202 participants, ensuring a strategic balance across gender, region, and age. Through detailed annotation across multiple dimensions, it can support a wide range of speech tasks. We perform extensive experiments on speaker verification, speaker diarization, speech recognition, and speech editing tasks, offering crucial insights for the development of speech technologies targeting this age group.

语音数据集老龄化中文语音说话人识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。