开源端到端情感语音大模型,支持低延迟情感回应生成。
OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model

- 基于流式交错解码实现低延迟语音生成。
- 自动化构建低成本高质量情感对话数据集。
- 适合研究情感交互、语音生成与开放源代码的学者。
情感交互是人机沟通的核心,需理解带有副语言特征的语音并生成情感丰富、表达自然的回应。然而当前最先进的情感语音语言模型日益封闭,其架构、数据与训练细节对研究者不透明。为推动情感语音语言模型的透明化研究,我们提出 OpenS2S——一个完全开源、透明且端到端的情感语音语言模型,支持高效情感语音交互。基于情感语音转写模型 BLSP-Emo,OpenS2S 采用流式交错解码架构,实现低延迟语音生成。为支持端到端训练,我们构建了自动化数据构造流程,利用大语言模型生成情感内容,并通过可控文本转语音系统引入说话人与情绪变化,以极低人力成本合成多样、高质量的情感语音对话。该模型包含数据集、模型权重及预训练、微调代码,全面开源,助力学术界加速情感语音系统创新。项目网页:https://casia-lm.github.io/OpenS2S
原文摘要 · Abstract (English)
Empathetic interaction is a cornerstone of human-machine communication, due to the need for understanding speech enriched with paralinguistic cues and generating emotional and expressive responses. However, the most powerful empathetic LSLMs are increasingly closed off, leaving the crucial details about the architecture, data and development opaque to researchers. Given the critical need for transparent research into the LSLMs and empathetic behavior, we present OpenS2S, a fully open-source, transparent and end-to-end LSLM designed to enable empathetic speech interactions. Based on our empathetic speech-to-text model BLSP-Emo, OpenS2S further employs a streaming interleaved decoding architecture to achieve low-latency speech generation. To facilitate end-to-end training, OpenS2S incorporates an automated data construction pipeline that synthesizes diverse, high-quality empathetic speech dialogues at low cost. By leveraging large language models to generate empathetic content and controllable text-to-speech systems to introduce speaker and emotional variation, we construct a scalable training corpus with rich paralinguistic diversity and minimal human supervision. We release the fully open-source OpenS2S model, including the dataset, model weights, pre-training and fine-tuning codes, to empower the broader research community and accelerate innovation in empathetic speech systems. The project webpage can be accessed at https://casia-lm.github.io/OpenS2S
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。