arXiv:2507.08768cs.SDcs.AI2025-07

测试老录音对语音识别的影响,发现说话人识别仍易受年龄和音道干扰。

On Barriers to Archival Audio Processing

  • 用联合国教科文组织的20世纪中期广播录音测试现有语音技术
  • 语音识别对多语口音处理较好,但说话人嵌入易受年龄和录音条件影响
  • 适合关注历史音频处理与语音系统鲁棒性的研究者

本研究利用联合国教科文组织独特的20世纪中期广播录音集,评估现代现成语音识别(LID)与说话人识别(SR)方法的鲁棒性,尤其关注多语言使用者和跨年龄录音的影响。结果表明,如Whisper等语音识别系统在处理第二语言及口音语音方面能力日益增强;然而,说话人嵌入仍是语音处理流程中的薄弱环节,易受声道、年龄和语言等因素引起的偏差影响。若档案机构希望使用说话人识别技术进行说话人索引,则需克服这些挑战。

原文摘要 · Abstract (English)

In this study, we leverage a unique UNESCO collection of mid-20th century radio recordings to probe the robustness of modern off-the-shelf language identification (LID) and speaker recognition (SR) methods, especially with respect to the impact of multilingual speakers and cross-age recordings. Our findings suggest that LID systems, such as Whisper, are increasingly adept at handling second-language and accented speech. However, speaker embeddings remain a fragile component of speech processing pipelines that is prone to biases related to the channel, age, and language. Issues which will need to be overcome should archives aim to employ SR methods for speaker indexing.

语音识别历史音频说话人识别鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。