arXiv:2506.08967cs.SDcs.CL2025-06被引 11

打造端到端语音问答模型,让机器直接说人话。

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

  • 用双码本音频编码器提取语言与语义特征
  • 1300亿参数模型+神经声码器生成高保真语音
  • 支持精准语音控制,适合语音交互系统开发

大型音频语言模型(LALMs)虽推动了人机智能交互发展,但依赖文本输出限制了其生成自然语音的能力,难以实现无缝音频交互。为此,我们提出Step-Audio-AQAA,一种用于音频查询-音频回答(AQAA)任务的全端到端大型音频语言模型。该模型采用双码本音频分词器提取语言与语义特征,配备1300亿参数主干大语言模型及神经声码器以实现高保真语音合成。通过交错输出文本与音频的后训练策略提升语义连贯性,并结合直接偏好优化(DPO)与模型合并技术提升性能。在StepEval-Audio-360基准测试中,Step-Audio-AQAA在语音控制等关键指标上超越现有先进LALMs。本工作为端到端LALMs提供了可行方案,凸显了基于标记的声码器在提升AQAA任务整体表现中的关键作用。

原文摘要 · Abstract (English)

Large Audio-Language Models (LALMs) have significantly advanced intelligent human-computer interaction, yet their reliance on text-based outputs limits their ability to generate natural speech responses directly, hindering seamless audio interactions. To address this, we introduce Step-Audio-AQAA, a fully end-to-end LALM designed for Audio Query-Audio Answer (AQAA) tasks. The model integrates a dual-codebook audio tokenizer for linguistic and semantic feature extraction, a 130-billion-parameter backbone LLM and a neural vocoder for high-fidelity speech synthesis. Our post-training approach employs interleaved token-output of text and audio to enhance semantic coherence and combines Direct Preference Optimization (DPO) with model merge to improve performance. Evaluations on the StepEval-Audio-360 benchmark demonstrate that Step-Audio-AQAA excels especially in speech control, outperforming the state-of-art LALMs in key areas. This work contributes a promising solution for end-to-end LALMs and highlights the critical role of token-based vocoder in enhancing overall performance for AQAA tasks.

音频生成大模型语音交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。