arXiv:2411.04699cs.CL2024-11中稿 · ACL被引 6

构建14种印度语大规模语音翻译数据集,提升真实场景下的翻译性能。

Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian Languages

  • 整合现有资源、网络爬取与合成数据,构建多语言高覆盖数据集。
  • 涵盖4.4万小时音频与1700万对齐文本段,支持跨语言语音翻译训练。
  • 开源模型与数据,助力学术界与产业界共建印度语语音翻译系统。

由于缺乏大规模、公开可用的数据集来捕捉印度语言的多样性和领域覆盖范围,印度语语音翻译仍面临巨大挑战。现有数据集仅涵盖部分印度语言,且难以支撑在真实场景中泛化的鲁棒模型训练。为此,我们提出了 BhasaAnuvaad,这是目前规模最大的印度语语音翻译数据集,涵盖超过44,000小时的音频和1700万条对齐文本段,覆盖14种印度语言及英语。该数据集通过三步法构建:(a) 整合高质量现有数据源,(b) 大规模网络爬取以确保语言和领域多样性,(c) 生成合成数据模拟真实语音不连贯性。基于 BhasaAnuvaad,我们训练了 IndicSeamless 模型,其在翻译质量上优于现有方法,为印度语语音翻译树立了新标准。我们将以宽松许可协议开源全部代码、数据与模型权重,推动开放协作。

原文摘要 · Abstract (English)

Speech translation for Indian languages remains a challenging task due to the scarcity of large-scale, publicly available datasets that capture the linguistic diversity and domain coverage essential for real-world applications. Existing datasets cover a fraction of Indian languages and lack the breadth needed to train robust models that generalize beyond curated benchmarks. To bridge this gap, we introduce BhasaAnuvaad, the largest speech translation dataset for Indian languages, spanning over 44 thousand hours of audio and 17 million aligned text segments across 14 Indian languages and English. Our dataset is built through a threefold methodology: (a) aggregating high-quality existing sources, (b) large-scale web crawling to ensure linguistic and domain diversity, and (c) creating synthetic data to model real-world speech disfluencies. Leveraging BhasaAnuvaad, we train IndicSeamless, a state-of-the-art speech translation model for Indian languages that performs better than existing models. Our experiments demonstrate improvements in the translation quality, setting a new standard for Indian language speech translation. We will release all the code, data and model weights in the open-source, with permissive licenses to promote accessibility and collaboration.

语音翻译多语言数据集印度语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。