arXiv:2509.09174cs.CLcs.AI2025-09被引 4

通过语音回声训练,提升语音大模型的推理能力

EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs

  • 用语义引导生成语音目标,融合声学与语义学习
  • 六千小时数据训练后,在多个问答任务上表现领先
  • 适合需要强推理的语音大模型研究者使用

语音到语音的大语言模型(SLLMs)正受到越来越多关注。源自文本型大语言模型(LLMs),SLLMs 常因知识与推理能力下降而受限。我们假设这一问题源于当前 SLLM 训练范式未能弥合特征空间中的声学-语义鸿沟。为此,我们提出 EchoX,利用语义表示并动态生成语音训练目标,实现声学与语义学习的融合,使 EchoX 在作为语音大模型时仍保持强推理能力。实验表明,使用约六千小时训练数据的 EchoX,在多个基于知识的问答基准测试中达到先进水平。项目代码已开源:https://github.com/FreedomIntelligence/EchoX。

原文摘要 · Abstract (English)

Speech-to-speech large language models (SLLMs) are attracting increasing attention. Derived from text-based large language models (LLMs), SLLMs often exhibit degradation in knowledge and reasoning capabilities. We hypothesize that this limitation arises because current training paradigms for SLLMs fail to bridge the acoustic-semantic gap in the feature representation space. To address this issue, we propose EchoX, which leverages semantic representations and dynamically generates speech training targets. This approach integrates both acoustic and semantic learning, enabling EchoX to preserve strong reasoning abilities as a speech LLM. Experimental results demonstrate that EchoX, with about six thousand hours of training data, achieves advanced performance on multiple knowledge-based question-answering benchmarks. The project is available at https://github.com/FreedomIntelligence/EchoX.

语音大模型声学语义推理增强语音生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。