arXiv:2506.04714cs.CLeess.AS2025-06中稿 · IWSLT2025被引 1

优化超参数与数据增强,提升低资源博杰普里语到印地语语音翻译性能

IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation

  • 系统调优学习率、训练步数等超参数,提升模型表现
  • 采用速度扰动和SpecAugment增强数据,显著改善翻译质量
  • 适合低资源语音翻译研究者参考,尤其关注小语种场景

本文介绍印度理工学院海得拉巴分校-布特大学(IIITH-BUT)在 IWSLT 2025 低资源博杰普里语-印地语语音翻译共享任务中的参赛方案。针对该语言对数据稀缺问题,我们对 SeamlessM4T 模型进行了微调,并系统研究了学习率调度、更新步数、预热步数、标签平滑及批量大小等超参数的影响。同时,为缓解数据不足,应用了速度扰动和 SpecAugment 数据增强技术,并评估其对翻译质量的提升效果。此外,还探索了通过联合训练马拉地语和博杰普里语语音数据引入跨语言信号的可行性。实验表明,在低资源条件下,精心选择超参数并结合简单有效的数据增强方法能显著提升翻译性能。我们进一步分析了翻译结果中的错误类型,揭示了影响 BLEU 分数的主要因素。

原文摘要 · Abstract (English)

This paper presents the submission of IIITH-BUT to the IWSLT 2025 shared task on speech translation for the low-resource Bhojpuri-Hindi language pair. We explored the impact of hyperparameter optimisation and data augmentation techniques on the performance of the SeamlessM4T model fine-tuned for this specific task. We systematically investigated a range of hyperparameters including learning rate schedules, number of update steps, warm-up steps, label smoothing, and batch sizes; and report their effect on translation quality. To address data scarcity, we applied speed perturbation and SpecAugment and studied their effect on translation quality. We also examined the use of cross-lingual signal through joint training with Marathi and Bhojpuri speech data. Our experiments reveal that careful selection of hyperparameters and the application of simple yet effective augmentation techniques significantly improve performance in low-resource settings. We also analysed the translation hypotheses to understand various kinds of errors that impacted the translation quality in terms of BLEU.

语音翻译低资源数据增强多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。