arXiv:2604.00613cs.CL2026-04

构建首个柯尔克语语音翻译数据集,解决拼写不统一问题提升翻译质量。

English to Central Kurdish Speech Translation: Corpus Creation, Evaluation, and Orthographic Standardization

  • 基于TED演讲构建9.1万句对的英-柯尔克语音译数据集
  • 拼写不统一使翻译准确率下降,标准化后性能提升3.0 BLEU
  • 适合低资源语言语音翻译研究者参考

我们提出KUTED,一个源自TED和TEDx演讲的中央柯尔克语语音到文本翻译(S2TT)数据集。该语料库包含91,000个句子对,涵盖170小时英语音频、165万词的英语文本和140万词的中央柯尔克语文本。我们在S2TT任务上评估KUTED,发现拼写变异会显著降低柯尔克语翻译性能,产生非标准输出。为此,我们提出一种系统性文本标准化方法,带来显著性能提升,实现更一致的翻译结果。在独立于TED演讲的测试集上,微调后的Seamless模型达到15.18 BLEU;在FLEURS基准上,相比Seamless基线提升3.0 BLEU。此外,我们还从头训练了一个Transformer模型,并评估了将Seamless(ASR)与NLLB(MT)结合的级联系统。

原文摘要 · Abstract (English)

We present KUTED, a speech-to-text translation (S2TT) dataset for Central Kurdish, derived from TED and TEDx talks. The corpus comprises 91,000 sentence pairs, including 170 hours of English audio, 1.65 million English tokens, and 1.40 million Central Kurdish tokens. We evaluate KUTED on the S2TT task and find that orthographic variation significantly degrades Kurdish translation performance, producing nonstandard outputs. To address this, we propose a systematic text standardization approach that yields substantial performance gains and more consistent translations. On a test set separated from TED talks, a fine-tuned Seamless model achieves 15.18 BLEU, and we improve Seamless baseline by 3.0 BLEU on the FLEURS benchmark. We also train a Transformer model from scratch and evaluate a cascaded system that combines Seamless (ASR) with NLLB (MT).

语音翻译低资源语言拼写标准化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。