arXiv:2502.17239cs.CLcs.SD2025-02被引 88

端到端语音交互模型,支持实时对话与问答。

Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

  • 文本引导语音生成,12.5Hz帧率多码本编码保留音色语义
  • 双阶段预训练保语言能力,音频头独立建模声学特征
  • 适合语音助手、实时对话系统研发者使用

我们提出Baichuan-Audio,一个端到端的音频大语言模型,无缝融合音频理解与生成。其采用文本引导的语音生成机制,实现具备理解与生成能力的实时语音交互。模型基于预训练语音识别(ASR)模型,以12.5 Hz帧率对语音进行多码本离散化,确保语音标记同时保留语义与声学信息。为增强建模效果,引入独立音频头处理音频标记,有效捕捉其独特特征。为缓解预训练中智能损失并保留原始大语言模型能力,提出两阶段预训练策略,在保持语言理解的同时提升音频建模性能。对齐后,模型在实时语音对话和问答任务中表现优异,展现强大泛化性与效率。代码、模型及训练数据已开源。

原文摘要 · Abstract (English)

We introduce Baichuan-Audio, an end-to-end audio large language model that seamlessly integrates audio understanding and generation. It features a text-guided aligned speech generation mechanism, enabling real-time speech interaction with both comprehension and generation capabilities. Baichuan-Audio leverages a pre-trained ASR model, followed by multi-codebook discretization of speech at a frame rate of 12.5 Hz. This multi-codebook setup ensures that speech tokens retain both semantic and acoustic information. To further enhance modeling, an independent audio head is employed to process audio tokens, effectively capturing their unique characteristics. To mitigate the loss of intelligence during pre-training and preserve the original capabilities of the LLM, we propose a two-stage pre-training strategy that maintains language understanding while enhancing audio modeling. Following alignment, the model excels in real-time speech-based conversation and exhibits outstanding question-answering capabilities, demonstrating its versatility and efficiency. The proposed model demonstrates superior performance in real-time spoken dialogue and exhibits strong question-answering abilities. Our code, model and training data are available at https://github.com/baichuan-inc/Baichuan-Audio

语音生成大模型端到端对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。