对比语音模型与大模型融合效果,发现直接集成更优。
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
- 构建首个系统性评测框架,对比6种语音大模型与16种主流方案。
- 多数语音大模型在噪声、长句等场景下超越传统分步系统。
- 语音基础模型单独使用效果差,需结合大语言模型才能高效翻译。
随着大语言模型(LLMs)拓展至非文本模态,语音作为原生输入的语音大模型(SpeechLLMs)应运而生,可直接处理语音并实现语音到文本翻译(ST)等任务,跳过传统转录流水线。然而,这种集成是否优于现有级联架构仍不明确。本文提出Hearing to Translate,首个全面评测套件,对6个先进SpeechLLMs与16种强基准系统(包括领先的语音基础模型与多语言大模型组合)进行对比。评测覆盖16个基准、13种语言对及9种挑战性条件,涵盖话语不连贯、噪音干扰、长篇语音等。结果显示,级联系统整体仍最可靠,但多数最新SpeechLLMs在多种场景下可达到甚至超过级联表现;而仅用语音基础模型则显著落后于二者,表明将大语言模型整合进模型或流程中是实现高质量语音翻译的关键。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) expand beyond text, integrating speech as a native modality has given rise to SpeechLLMs, which directly process spoken language and enable speech-to-text translation (ST) and other downstream tasks, bypassing traditional transcription-based pipelines. Whether this integration improves ST quality over established cascaded architectures, however, remains an open question. We present Hearing to Translate, the first comprehensive test suite rigorously benchmarking 6 state-of-the-art SpeechLLMs against 16 strong direct and cascade systems that couple leading speech foundation models (SFM), with multilingual LLMs. Our analysis spans 16 benchmarks, 13 language pairs, and 9 challenging conditions, including disfluent, noisy, and long-form speech. Across this extensive evaluation, we find that cascaded systems remain the most reliable solution overall, but most recent SpeechLLMs can match or even outperform cascades in various settings while SFMs lag behind both, highlighting that integrating an LLM, either within the model or in a pipeline, is essential for high-quality speech translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。