arXiv:2502.04537cs.CL2025-02被引 2

无需知识蒸馏的多语言非自回归翻译新方法

Multilingual Non-Autoregressive Machine Translation without Knowledge Distillation

  • 基于有向无环Transformer结构,避免传统知识蒸馏
  • 在多语言翻译任务上达到当前最佳性能
  • 适合追求高效部署且需支持未见方向的场景

多语言神经机器翻译(MNMT)旨在使用单一模型实现多种翻译方向。近期工作采用非自回归Transformer提升效率,但需耗费高昂的知识蒸馏(KD)过程。为此,我们提出M-DAT方法,用于非自回归多语言机器翻译。系统利用最近提出的有向无环Transformer(DAT)结构,无需知识蒸馏。此外,我们提出一种枢纽回译(PivotBT)策略,以增强对未见翻译方向的泛化能力。实验表明,M-DAT在非自回归多语言机器翻译中达到了当前最优性能。

原文摘要 · Abstract (English)

Multilingual neural machine translation (MNMT) aims at using one single model for multiple translation directions. Recent work applies non-autoregressive Transformers to improve the efficiency of MNMT, but requires expensive knowledge distillation (KD) processes. To this end, we propose an M-DAT approach to non-autoregressive multilingual machine translation. Our system leverages the recent advance of the directed acyclic Transformer (DAT), which does not require KD. We further propose a pivot back-translation (PivotBT) approach to improve the generalization to unseen translation directions. Experiments show that our M-DAT achieves state-of-the-art performance in non-autoregressive MNMT.

多语言翻译非自回归Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。