arXiv:2505.13404cs.CLeess.AS2025-05中稿 · Interspeech 2025 v…被引 30

构建25种欧洲语言的语音识别与翻译数据集,解决低资源语言数据匮乏问题。

Granary: Speech Recognition and Translation Dataset in 25 European Languages

  • 用伪标签+两轮推理提升语音数据质量,自动过滤幻觉并恢复标点。
  • 基于伪标签生成翻译对,仅用一半数据达到现有模型同等性能。
  • 首个开源大规模多语言语音数据集,适合语音与跨语言研究者使用。

多任务与多语言方法使大模型受益,但低资源语言的语音处理仍因数据稀缺而受限。为此,我们提出Granary,一个涵盖25种欧洲语言的语音识别与翻译大规模数据集,是首个在转录与翻译两方面均达到此规模的开源项目。通过分割、双轮推理、幻觉过滤与标点恢复的伪标签流水线提升数据质量,并利用EuroLLM从伪标签转录中生成翻译对,再经数据过滤流程优化。该流水线高效,可在数小时内处理海量数据。我们在已有的高/低资源语言数据集上评估训练模型,结果表明,使用约50%更少的数据即可达到相近性能。数据集将公开于https://hf.co/datasets/nvidia/Granary。

原文摘要 · Abstract (English)

Multi-task and multilingual approaches benefit large models, yet speech processing for low-resource languages remains underexplored due to data scarcity. To address this, we present Granary, a large-scale collection of speech datasets for recognition and translation across 25 European languages. This is the first open-source effort at this scale for both transcription and translation. We enhance data quality using a pseudo-labeling pipeline with segmentation, two-pass inference, hallucination filtering, and punctuation restoration. We further generate translation pairs from pseudo-labeled transcriptions using EuroLLM, followed by a data filtration pipeline. Designed for efficiency, our pipeline processes vast amount of data within hours. We assess models trained on processed data by comparing their performance on previously curated datasets for both high- and low-resource languages. Our findings show that these models achieve similar performance using approx. 50% less data. Dataset will be made available at https://hf.co/datasets/nvidia/Granary

语音识别多语言数据集低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。