用语音嵌入选数据,5%精选样本胜过全量训练
Which Data Matter? Embedding-Based Data Selection for Speech Recognition
- 基于说话人、发音和语义三维度嵌入,筛选相关数据
- 仅用5%精选数据,目标领域字错率降低36.8%相对值
- 适合专注特定场景的语音识别模型优化
现代语音识别系统通常在大规模伪标注、跨领域的野外数据上训练。虽然这类异构数据有助于通用模型,但对针对特定领域的专用模型带来挑战:专用模型难以处理全部数据,且需关注训练与测试条件的不匹配问题。本文研究针对性数据选择策略,从10万小时野外数据中选取子集以优化目标领域性能。通过使用捕捉说话人属性、发音内容和语义意义的语音嵌入,分析不同维度的相关性与多样性对下游语音识别的影响。实验采用基于CTC的Conformer模型,结果表明,仅用经策略性选择的5%数据训练,即可在目标领域实现最高达36.8%相对字错率降低,优于全量数据训练模型。
原文摘要 · Abstract (English)
Modern ASR systems are typically trained on large-scale pseudo-labeled, in-the-wild data spanning multiple domains. While such heterogeneous data benefit generalist models designed for broad deployment, they pose challenges for specialist models targeting specific domains: specialist models lack the capacity to learn from all available data, and one must pay closer attention to addressing the mismatch between training and test conditions. In this work, we study targeted data selection as a strategy to address these challenges, selecting relevant subsets from 100k hours of in-the-wild training data to optimize performance on target domains. We represent speech samples using embeddings that capture complementary characteristic--speaker attributes, phonetic content, and semantic meaning--and analyze how relevance and diversity along these axes when performing data selection affect downstream ASR performance. Our experiments with CTC-based Conformer models show that training on a strategically selected 5% subset can exceed the performance of models trained on the full dataset by up to 36.8% relative WER reduction on target domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。