arXiv:2511.14410eess.AS2025-11中稿 · ICASSP2026被引 3

轻量级模型TTA提升跨语言语音表征,更好对接大语言模型。

TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation

  • 基于35.8万小时多语种语音数据,联合训练语音识别、翻译与对齐任务。
  • 在跨语言语音理解等任务上超越Whisper,尤其在多语种场景表现更优。
  • 适合需要高效跨语言语音处理的研究者和开发者使用。

语音-大语言模型(Speech-LLM)在多模态、多任务语音理解中表现优异。典型范式是将语音模态与大语言模型(LLM)融合。尽管先前研究常采用Whisper编码器作为语音输入,但其在输入格式、模型规模和语义性能方面存在局限。为此,我们提出一种专用于语音语义的轻量级模型TTA,以更有效地实现与LLM的集成。通过在35.8万小时多语种语音数据上进行大规模训练,涵盖语音识别(ASR)、语音翻译(ST)和语音-文本对齐任务,TTA能够生成稳健的跨语言语音表征。在多个基准测试中,包括ASR/ST、语音检索以及ASR-LLM性能评估,TTA均优于Whisper。此外,我们严格验证了跨语言能力与ASR/ST性能之间的相互作用。TTA的模型权重与训练方案将作为音频理解工具包Auden的一部分发布。

原文摘要 · Abstract (English)

Speech-LLM models have demonstrated great performance in multi-modal and multi-task speech understanding. A typical speech-LLM paradigm is integrating speech modality with a large language model (LLM). While the Whisper encoder was frequently adopted in previous studies for speech input, it shows limitations regarding input format, model scale, and semantic performance. To this end, we propose a lightweight TTA model specialized in speech semantics for more effective LLM integration. With large-scale training of 358k hours of speech data on multilingual speech recognition (ASR), speech translation (ST) and speech-text alignment tasks, TTA is capable of producing robust cross-lingual speech representations. Extensive evaluations across diverse benchmarks, including ASR/ST, speech retrieval, and ASR-LLM performance assessments, demonstrate TTA's superiority over Whisper. Furthermore, we rigorously validate the interplay between cross-lingual capabilities and ASR/ST performance. The model weights and training recipes of TTA will be released as part of an audio understanding toolkit Auden.

语音表征跨语言大模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。