增加说话人数量比延长单人录音时长更提升低资源语音识别抗方言能力
Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR
- 在总训练时长固定时,多说话人比单人长录音更有利
- 说话人越多,模型性能随训练时长增长越明显
- 控制说话人数量时,方言多样性影响微弱,优先扩增人数
为构建能服务全球用户的自动语音识别系统,ASR需具备对未见方言的鲁棒性。本文系统研究了训练数据中三个变量——说话人数量、每位说话人的音频时长、方言多样性——在低资源训练条件下对未见方言鲁棒性的影响。在固定总训练时长下,增加说话人数量(即每人贡献较少)比增加单人时长更有效。同时发现,更多说话人能使模型性能随训练时长提升更显著。令人意外的是,在说话人数量受限时,优先选择不同方言说话人带来的收益微乎其微。研究建议,新语言的ASR训练应优先增加说话人数量。
原文摘要 · Abstract (English)
To build an automatic speech recognition (ASR) system that can serve everyone in the world, the ASR needs to be robust to a wide range of accents including unseen accents. We systematically study how three different variables in training data -- the number of speakers, the audio duration per each individual speaker, and the diversity of accents -- affect ASR robustness towards unseen accents in a low-resource training regime. We observe that for a fixed number of ASR training hours, it is more beneficial to increase the number of speakers (which means each speaker contributes less) than the number of hours contributed per speaker. We also observe that more speakers enables ASR performance gains from scaling number of hours. Surprisingly, we observe minimal benefits to prioritizing speakers with different accents when the number of speakers is controlled. Our work suggests that practitioners should prioritize increasing the speaker count in ASR training data composition for new languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。