针对低资源语言对话,构建并优化了语音到语音翻译流水线。
Speech-to-Speech Translation Pipelines for Conversations in Low-Resource Languages
- 分步构建语音识别、机器翻译、语音合成流水线,支持本地与云端模型。
- 在土耳其语/普什图语与法语互译中,测试超60种组合,找到最优方案。
- 组件表现独立于整体流程,可灵活替换,适合社区口译场景。
语音到语音翻译在人际对话中的应用日益广泛,但翻译质量因语言对而异。针对土耳其语和普什图语与法语之间的社区口译需求,我们收集了微调和测试数据,通过多种自动评估指标(BLEU、COMET、BLASER)及人工评测,对比了包含自动语音识别、机器翻译和语音合成的多个系统。系统采用本地模型与云端商业模型混合配置,部分组件在我们的数据上进行了微调。共评估了超过60种流水线配置,确定了双向最优方案。研究还发现,各组件的表现排名在很大程度上独立于流水线其他部分。
原文摘要 · Abstract (English)
The popularity of automatic speech-to-speech translation for human conversations is growing, but the quality varies significantly depending on the language pair. In a context of community interpreting for low-resource languages, namely Turkish and Pashto to/from French, we collected fine-tuning and testing data, and compared systems using several automatic metrics (BLEU, COMET, and BLASER) and human assessments. The pipelines included automatic speech recognition, machine translation, and speech synthesis, with local models and cloud-based commercial ones. Some components have been fine-tuned on our data. We evaluated over 60 pipelines and determined the best one for each direction. We also found that the ranks of components are generally independent of the rest of the pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。