arXiv:2605.23912cs.CLcs.AI2026-05

90亿参数语音模型,支持英韩双语听懂、回答和生成。

Raon-Speech Technical Report

论文配图:Raon-Speech Technical Report
图 1 · 摘自论文原文
  • 用三阶段训练将文本大模型转为可语音输入输出的通用语音模型。
  • 在42个基准上超越8个同类模型,语音任务表现最强。
  • 开源全链路工具,适合语音交互与对话系统研发者使用。

我们提出Raon-Speech,一个90亿参数的英语与韩语语音理解、问答与生成的高性能语音语言模型(SpeechLM),以及其全双工扩展Raon-SpeechChat,用于自然实时对话。Raon-Speech成功将预训练文本大模型转化为兼具语音理解与生成能力且保留强文本性能的SpeechLM。其训练基于138万小时经严格筛选的英韩语音与文本数据,包含三个阶段:(1) 语音模块对齐,(2) 基于知识蒸馏的端到端SpeechLM预训练,(3) 多任务偏好优化后训练。在42个英韩语音与文本基准中,相比八款同规模近期音频基础模型(包括Qwen2.5-Omni与Fun-Audio-Chat),Raon-Speech在语音核心任务上表现最强,同时保持优异文本问答能力。在此基础上,Raon-SpeechChat通过持续训练11.9万小时时间对齐的真实与合成对话数据,分三阶段实现:(1) 因果编码器适配,(2) 全双工预训练,(3) 语音与角色控制的全双工微调。在多个全双工基准测试中,其在FDB v1.0涵盖的轮次切换与打断敏感行为上优势明显,整体表现仍具竞争力。所有模型检查点、训练与推理流程及交互式演示均已开源。

原文摘要 · Abstract (English)

We present Raon-Speech, a top-performing 9B-parameter speech language model (SpeechLM) for English and Korean speech understanding, answering, and generation, and Raon-SpeechChat, a high-performing full-duplex extension for natural real-time conversation. Raon-Speech successfully transforms a pre-trained LLM into a SpeechLM that both understands and generates speech while preserving strong text capabilities. It trains on 1.38M hours of highly curated English and Korean speech and text datasets with the following training stages: (1) speech modules alignment, (2) end-to-end SpeechLM pre-training with knowledge distillation, and (3) multi-task preference optimization-based post-training. Across 42 English and Korean speech and text benchmarks, Raon-Speech establishes the strongest overall profile on speech-centric tasks in our comparison against eight similarly sized recent audio foundation models, including Qwen2.5-Omni and Fun-Audio-Chat, while preserving strong text question answering performance. Building upon it, Raon-SpeechChat enables natural full-duplex conversation by continual training on 119K hours of time-aligned real and synthetic dialogue data. It proceeds through three complementary training stages: (1) causal encoder adaptation, (2) full-duplex pre-training, (3) full-duplex fine-tuning for voice and role-control. On multiple full-duplex benchmarks, Raon-SpeechChat shows its clearest strengths on the turn-taking and interruption-sensitive behaviors covered by FDB v1.0, and remains competitive across the broader full-duplex evaluation suite. We open-source all model checkpoints, the training and inference pipeline, and an interactive demo.

语音模型全双工对话多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。