arXiv:2503.10211cs.CLcs.SD2025-03被引 3

通过动态对齐语音与文本表示,提升大模型语音翻译效果

Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation

  • 在大模型特定层显式对齐语音与文本表征
  • 利用最优传输理论量化跨模态差异,显著提升翻译准确率
  • 适合研究跨模态对齐与语音翻译的开发者

大语言模型(LLMs)的进展为语音翻译系统的发展奠定了基础。现有方法主要关注跨模态输入输出对齐,却忽略了模型内部表示层面的深层语义对齐。为此,本文提出自适应内层语音-文本对齐(AI-STA)方法,通过在大模型选定层显式对齐语音与文本表征来弥合模态差距。我们采用最优传输(OT)理论量化语音与文本表示间的细粒度差异,并结合跨模态检索技术识别最适合对齐的层,对这些层进行联合训练。在语音翻译任务上的实验结果表明,AI-STA显著提升了大语音-文本模型(LSMs)的翻译性能,优于此前最先进方法。研究揭示了大模型内部层间对齐的重要性,为增强跨模态学习提供了新思路。

原文摘要 · Abstract (English)

Recent advancement of large language models (LLMs) has led to significant breakthroughs across various tasks, laying the foundation for the development of LLM-based speech translation systems. Existing methods primarily focus on aligning inputs and outputs across modalities while overlooking deeper semantic alignment within model representations. To address this limitation, we propose an Adaptive Inner Speech-Text Alignment (AI-STA) method to bridge the modality gap by explicitly aligning speech and text representations at selected layers within LLMs. To achieve this, we leverage the optimal transport (OT) theory to quantify fine-grained representation discrepancies between speech and text. Furthermore, we utilize the cross-modal retrieval technique to identify the layers that are best suited for alignment and perform joint training on these layers. Experimental results on speech translation (ST) tasks demonstrate that AI-STA significantly improves the translation performance of large speech-text models (LSMs), outperforming previous state-of-the-art approaches. Our findings highlight the importance of inner-layer speech-text alignment in LLMs and provide new insights into enhancing cross-modal learning.

语音翻译大模型跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。