arXiv:2506.00733eess.AScs.SD2025-06中稿 · Interspeech 2025被引 5

用声音嵌入量化并减少语音数据中的说话人混杂问题

Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis

  • 通过ResNet提取声音特征,计算同客户端录音的相似度
  • 设定最优阈值后,使同一客户端内仅保留高度相似录音
  • 提升语音分析与说话人相关技术的数据质量

Mozilla Common Voice语料库(CV)因其跨语言、跨说话人的多样性,成为多语言语音技术的重要资源,并在跨语言语音学和语音科学中具有巨大潜力。然而,准确处理说话人差异是语音研究理论与统计基础的关键。尽管CV提供客户端ID作为说话人标识的近似,但同一客户端可能由多个说话人贡献。本研究旨在量化并减少客户端ID内的说话人异质性,以更接近真实但匿名的说话人身份。我们使用基于ResNet的声音嵌入,计算相同客户端下录音间的相似度,并通过说话人判别任务确定最优阈值,从而降低感知到的说话人异质性。该结果对语音学分析及基于说话人的语音技术开发具有重要应用价值。

原文摘要 · Abstract (English)

With its crosslinguistic and cross-speaker diversity, the Mozilla Common Voice Corpus (CV) has been a valuable resource for multilingual speech technology and holds tremendous potential for research in crosslinguistic phonetics and speech sciences. Properly accounting for speaker variation is, however, key to the theoretical and statistical bases of speech research. While CV provides a client ID as an approximation to a speaker ID, multiple speakers can contribute under the same ID. This study aims to quantify and reduce heterogeneity in the client ID for a better approximation of a true, though still anonymous speaker ID. Using ResNet-based voice embeddings, we obtained a similarity score among recordings with the same client ID, then implemented a speaker discrimination task to identify an optimal threshold for reducing perceived speaker heterogeneity. These results have major downstream applications for phonetic analysis and the development of speaker-based speech technology.

语音分析说话人识别数据清洗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。