构建首个希伯来语跨龄语音数据集,助力语音系统应对声音老化挑战
VoxKnesset: A Large-Scale Longitudinal Hebrew Speech Dataset for Aging Speaker Modeling
- 采集2009至2025年议会发言,覆盖393人,最长跨度15年
- 模型在15年内说话人验证错误率从2.15%升至4.58%,揭示老化影响
- 首次提供纵向训练框架,适合语音老化研究与希伯来语系统开发
语音处理系统面临核心挑战:人类嗓音随年龄变化,但缺乏支持严格纵向评估的数据集。本文提出VoxKnesset,一个开放获取的希伯来语议会发言数据集,包含约2300小时语音,覆盖393名发言人,时间跨度达15年(2009–2025)。每段音频均配有对齐转录和来自官方记录的经核实人口统计学元数据。我们以现代语音嵌入模型(WavLM-Large、ECAPA-TDNN、Wav2Vec2-XLSR-1B)在纵向条件下评估年龄预测与说话人验证性能。结果显示,最强模型在15年间说话人验证等错误率(EER)从2.15%升至4.58%;而横截面训练的年龄回归模型无法捕捉个体老化趋势,纵向训练模型则能恢复显著的时间信号。数据集与处理管道已公开发布,旨在推动抗老化语音系统与希伯来语语音技术发展。
原文摘要 · Abstract (English)
Speech processing systems face a fundamental challenge: the human voice changes with age, yet few datasets support rigorous longitudinal evaluation. We introduce VoxKnesset, an open-access dataset of ~2,300 hours of Hebrew parliamentary speech spanning 2009-2025, comprising 393 speakers with recording spans of up to 15 years. Each segment includes aligned transcripts and verified demographic metadata from official parliamentary records. We benchmark modern speech embeddings (WavLM-Large, ECAPA-TDNN, Wav2Vec2-XLSR-1B) on age prediction and speaker verification under longitudinal conditions. Speaker verification EER rises from 2.15\% to 4.58\% over 15 years for the strongest model, and cross-sectionally trained age regressors fail to capture within-speaker aging, while longitudinally trained models recover a meaningful temporal signal. We publicly release the dataset and pipeline to support aging-robust speech systems and Hebrew speech processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。