arXiv:2504.07024cs.CL2025-04被引 1

小语种语音识别精度低?调参比数据增强更有效。

Data Augmentation and Hyperparameter Tuning for Low-Resource MFA

  • 用超参数调优替代数据增强,提升小语种语音对齐效果。
  • 在小到中等数据量下,调参使准确率显著提升。
  • 适合资源匮乏语言的语音分析研究者使用。

处理濒危或资源匮乏语言时,计算工具在数据量少的语言上表现较差。本文通过数据增强扩大语料库规模,比较了音频增强与超参数调优在多语言强制对齐中的效果。与文本增强不同,音频增强并未带来明显性能提升;而超参数调优在不增加训练时间的前提下实现了显著改进。对于小到中等规模训练数据的语言,该方法是无需从高资源语言迁移模型的有效替代方案。

原文摘要 · Abstract (English)

A continued issue for those working with computational tools and endangered and under-resourced languages is the lower accuracy of results for languages with smaller amounts of data. We attempt to ameliorate this issue by using data augmentation methods to increase corpus size, comparing augmentation to hyperparameter tuning for multilingual forced alignment. Unlike text augmentation methods, audio augmentation does not lead to substantially increased performance. Hyperparameter tuning, on the other hand, results in substantial improvement without (for this amount of data) infeasible additional training time. For languages with small to medium amounts of training data, this is a workable alternative to adapting models from high-resource languages.

语音识别小样本调参

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。