arXiv:2504.16234cs.LGcs.CL2025-04中稿 · Swiss NLP Conferen…被引 1

用音素替代文字,提升多语种语音翻译效率

Using Phonemes in cascaded S2S translation pipeline

  • 在语音翻译流水线中用音素代替文本作为表示
  • 音素方法达到与文本相当的翻译质量(BLEU)
  • 更适合资源匮乏语言,降低计算资源需求

本文探索在传统多语种同步语音到语音翻译管道中使用音素作为文本表示的可行性,而非依赖传统的文本语言表示。为此,我们在WMT17数据集上训练了一个开源序列到序列模型,分别采用标准文本表示和音素表示两种格式。通过BLEU指标评估两者性能。结果表明,音素方法在翻译质量上与文本方法相当,但具有更低的资源需求或对低资源语言更友好。

原文摘要 · Abstract (English)

This paper explores the idea of using phonemes as a textual representation within a conventional multilingual simultaneous speech-to-speech translation pipeline, as opposed to the traditional reliance on text-based language representations. To investigate this, we trained an open-source sequence-to-sequence model on the WMT17 dataset in two formats: one using standard textual representation and the other employing phonemic representation. The performance of both approaches was assessed using the BLEU metric. Our findings shows that the phonemic approach provides comparable quality but offers several advantages, including lower resource requirements or better suitability for low-resource languages.

语音翻译音素表示低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。