通过最优传输对齐跨语言语音,提升低资源语种翻译效果
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
- 用最优传输技术对齐多语言语音表征,减少翻译偏差
- 在五种常用语言上提升1.29 BLEU,零样本语言提升2.93 BLEU
- 仅需每语言10小时平行语音数据,适合低资源场景
语音大模型在多语言语音转文本翻译中取得突破,但现有方法常忽视源语言间的语义共性,导致翻译性能偏倚。本文提出POTSA(基于并行最优传输的语音对齐框架),利用跨语言并行语音对与最优传输技术,弥合高、低资源翻译差距。首先引入偏差补偿模块粗对齐初始语音表示;其次在Q-Former上施加词级别最优传输约束,建立细粒度表示一致性;最后采用层调度策略,将最优传输约束聚焦于语义有益层。在FLEURS数据集上的实验表明,本方法达到当前最优性能,五种常见语言平均提升1.29 BLEU,零样本语言提升2.93 BLEU,且每语言仅需10小时平行语音数据。
原文摘要 · Abstract (English)
Speech Large Language Models have achieved breakthroughs in multilingual speech-to-text translation. However, existing approaches often overlook semantic commonalities across source languages, leading to biased translation performance. In this work, we propose POTSA (Parallel Optimal Transport for Speech Alignment), a new framework based on cross-lingual parallel speech pairs and Optimal Transport, designed to bridge high- and low-resource translation gaps. First, we introduce a Bias Compensation module to coarsely align initial speech representations. Second, we impose token-level OT constraints on a Q-Former using parallel pairs to establish fine-grained representation consistency. Then, we apply a layer scheduling strategy to focus OT constraints on semantically beneficial layers. Experiments on FLEURS show our method achieves SOTA performance, with +1.29 BLEU over five common languages and +2.93 BLEU on zero-shot languages, using only 10 hours of parallel speech per language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。