用字符级建模提升低资源语言翻译能力,支持上千种语言的语音翻译。
Improving Language and Modality Transfer in Translation by Character-level Modeling
- 基于字符级编码器与SONAR嵌入空间,实现跨语言跨模态知识迁移。
- 在75种语言文本翻译中优于传统子词模型,低资源语言表现显著提升。
- 零样本适配语音翻译,仅需少量ASR数据即达领先性能,适合多语言系统开发。
当前翻译系统虽具备高度多语言能力,但仅覆盖全球语言的5%。扩展至长尾低资源语言需数据高效方法,依赖跨语言与跨模态知识迁移。为此,我们提出一种字符级建模方法,以提升对新语言和模态的适应性。该方法利用SONAR——一个具有编码与解码模块的多语言固定尺寸嵌入空间。通过并行翻译数据,采用教师-学生框架训练字符级编码器;再结合语音识别(ASR)数据,训练轻量适配器,将大规模多语言CTC语音识别模型(MMS)连接至字符级编码器,有望实现1000+语言的语音翻译。在FLORES+上对75种语言的文本翻译实验表明,该字符级方法在语言迁移方面优于传统子词模型,尤其在低资源设置下表现更优,并展现出更强的零样本泛化能力。语音适配部分最大化从文本模态的知识迁移,在FLEURS基准上33种语言的语音转文字翻译中达到当前最优结果,超越先前的监督与级联模型,且仅为零样本模型,仅需极少量来自ASR的数据进行微调。
原文摘要 · Abstract (English)
Current translation systems, despite being highly multilingual, cover only 5% of the world's languages. Expanding language coverage to the long-tail of low-resource languages requires data-efficient methods that rely on cross-lingual and cross-modal knowledge transfer. To this end, we propose a character-based approach to improve adaptability to new languages and modalities. Our method leverages SONAR, a multilingual fixed-size embedding space with different modules for encoding and decoding. We use a teacher-student approach with parallel translation data to obtain a character-level encoder. Then, using ASR data, we train a lightweight adapter to connect a massively multilingual CTC ASR model (MMS), to the character-level encoder, potentially enabling speech translation from 1,000+ languages. Experimental results in text translation for 75 languages on FLORES+ demonstrate that our character-based approach can achieve better language transfer than traditional subword-based models, especially outperforming them in low-resource settings, and demonstrating better zero-shot generalizability to unseen languages. Our speech adaptation, maximizing knowledge transfer from the text modality, achieves state-of-the-art results in speech-to-text translation on the FLEURS benchmark on 33 languages, surpassing previous supervised and cascade models, albeit being a zero-shot model with minimal supervision from ASR data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。