arXiv:2511.10670cs.CLcs.AI2025-11中稿 · IJCAI被引 1

通过语义空间对齐提升多语言语音翻译精度

Towards Fine-Grained Code-Switch Speech Translation with Semantic Space Alignment

  • 用专家混合模型分离不同语言的语义空间,实现细粒度特征建模
  • 在多个数据集上比SeamlessM4T提升最高1.49 BLEU和1.41 COMET
  • 适合需要高精度多语言语音翻译的场景,如跨语言客服系统

多语言切换(CS)语音翻译旨在将混合多种语言的语音翻译为单一目标语言文本,因语义建模复杂且数据稀缺而面临挑战。以往方法依赖模型隐式学习语义表示,需大量人工标注。为此,本文提出在大语言模型中引入由语言专家组构成的混合专家(MoE)语音投影器,每个专家组专注于特定语言的语义空间,实现细粒度语音特征建模。同时引入语言特异性损失与组内负载均衡损失,优化专家间与组内令牌路由。进一步设计多阶段训练范式,利用现成的自动语音识别(ASR)与单语种语音翻译数据,促进语音-文本对齐并提升翻译性能。为缓解领域迁移的数据鸿沟,引入过渡损失以增强对多语言切换场景的适应性。在多个常用数据集上的实验表明,该方法有效且具有通用性,平均提升0.86 BLEU与0.93 COMET,最大提升达1.49 BLEU与1.41 COMET。

原文摘要 · Abstract (English)

Code-switching (CS) speech translation (ST) aims to translate speech that alternates between multiple languages into a target language text, posing significant challenges due to the complexity of semantic modeling and the scarcity of CS data. Previous studies mainly rely on the models themselves to implicitly learn semantic representations and resort to costly manual annotations. To mitigate these limitations, we propose enhancing Large Language Models (LLMs) with a Mixture-of-Experts (MoE) speech projector composed of language expert groups, where each group specializes in the semantic space of a specific language for fine-grained speech feature modeling. A language-specific loss and an intra-group load balancing loss are jointly introduced to guide efficient token routing across and within expert groups. Furthermore, we introduce a multi-stage training paradigm that utilizes readily available automatic speech recognition (ASR) and monolingual ST data, facilitating speech-text alignment and improving translation performance. To bridge the data gap for smooth domain transfer, a transition loss is employed to improve adaptation to CS scenarios. Extensive experiments on widely used datasets demonstrate the effectiveness and generality of our approach, achieving average improvements of $0.86$ BLEU and $0.93$ COMET over SeamlessM4T, with maximum improvements of $1.49$ BLEU and $1.41$ COMET across different test sets.

语音翻译多语言MoE语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。