arXiv:2603.25836cs.CL2026-03被引 1

通过分析梯度自动优化多语言语音翻译的共享结构,提升低资源场景表现。

Gradient-Informed Training for Low-Resource Multilingual Speech Translation

  • 根据梯度信息动态确定每层的语言共享模式
  • 在四种语言对上显著提升翻译质量
  • 适合低资源多语言语音翻译研究者参考

在低资源多语言语音到文本翻译中,跨语言统一架构常引发表示冲突,阻碍模型收敛。本文提出一种基于梯度信息的系统性方法,自动识别层级共享模式。通过三种分析策略:基于距离的语言聚类、自任务/跨任务发散度度量以分配容量,以及联合分解结合典型相关分析实现子空间对齐。在四个语言对(采用SeamlessM4T-Medium架构)上进行充分评估,持续提升翻译质量指标。

原文摘要 · Abstract (English)

In low-resource multilingual speech-to-text translation, uniform architectural sharing across languages frequently introduces representation conflicts that impede convergence. This work proposes a principled methodology to automatically determine layer-specific sharing patterns by mining training gradient information. Our approach employs three distinct analysis strategies: distance-based language clustering, self/cross-task divergence metrics for capacity allocation, and joint factorization coupled with canonical correlation analysis for subspace alignment. Extensive evaluation across four language pairs (using the SeamlessM4T-Medium architecture) demonstrates persistent improvements in translation quality metrics.

多语言语音翻译低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。