arXiv:2510.00499cs.CLcs.AI2025-10被引 10

无需文本中间步骤,直接实现语音到语音的自然对话。

MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance

论文配图:MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance
图 1 · 摘自论文原文
  • 采用分层模态架构+冻结预训练策略,融合文本大模型知识与语音生成能力。
  • 在语音问答任务上达到顶尖水平,语音生成质量媲美依赖文本的系统。
  • 适合追求自然表达、低延迟语音交互的研究者与开发者。

语音对话系统通常依赖转录、处理、重合成的级联流程,虽有效但会丢失语气等副语言特征并限制表达力。近期端到端方法虽降低延迟并更好保留这些特征,但仍依赖文本中间表示,存在根本性瓶颈。本文提出 MOSS-Speech,一种真正无需文本引导的语音到语音大型语言模型,可直接理解与生成语音。方法结合基于模态的层分裂架构与冻结预训练策略,在保留预训练文本大模型推理与知识能力的同时,新增原生语音处理能力。实验表明,该模型在语音问答任务上达到当前最优性能,语音到语音生成表现可与现有文本引导系统相媲美,同时保持竞争力的文本任务表现。通过缩小文本引导与直接语音生成之间的差距,本工作为更自然、高效的端到端语音交互树立了新范式。

原文摘要 · Abstract (English)

Spoken dialogue systems often rely on cascaded pipelines that transcribe, process, and resynthesize speech. While effective, this design discards paralinguistic cues and limits expressivity. Recent end-to-end methods reduce latency and better preserve these cues, yet still rely on text intermediates, creating a fundamental bottleneck. We present MOSS-Speech, a true speech-to-speech large language model that directly understands and generates speech without relying on text guidance. Our approach combines a modality-based layer-splitting architecture with a frozen pre-training strategy, preserving the reasoning and knowledge of pretrained text LLMs while adding native speech capabilities. Experiments show that our model achieves state-of-the-art results in spoken question answering and delivers comparable speech-to-speech performance relative to existing text-guided systems, while still maintaining competitive text performance. By narrowing the gap between text-guided and direct speech generation, our work establishes a new paradigm for expressive and efficient end-to-end speech interaction.

语音生成大模型端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。