首个支持角色扮演与唱歌的端到端语音模型,让对话更生动有表现力。
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing

- 采用文本音频混合建模与多码本音频令牌,分离模态又增强表达
- 训练数据达15.8小时,角色扮演评测领先同行7个百分点,唱歌质量提升0.13分
- 适合需要情感化语音生成的应用,如虚拟角色、音乐互动、智能客服
人类语音不仅传递语言内容,还包含个性、情绪或表演元素,如温柔语调或哼唱歌曲,我们将其形式化为角色扮演与歌唱。本文提出VITA-QinYu,首个支持角色扮演与歌唱生成的端到端语音语言模型。该模型采用混合文本-音频范式,通过多码本音频令牌扩展交错文本-音频建模,增强副语言表征能力,同时保持模态清晰分离,避免干扰。我们构建了完整的数据生成流程,合成总计15.8小时的自然对话、角色扮演与歌唱数据用于训练。VITA-QinYu在客观角色扮演评测中领先同行7个百分点,在5分制主观评分(MOS)中歌唱表现超出对手0.13分;同时在对话准确性和流畅性上达到业界领先水平,分别在C3和URO基准上超越前序模型1.38和4.98个百分点。项目已开源代码与模型,并提供支持流式与全双工交互的演示系统。
原文摘要 · Abstract (English)
Human speech conveys expressiveness beyond linguistic content, including personality, mood, or performance elements, such as a comforting tone or humming a song, which we formalize as role-playing and singing. We present VITA-QinYu, the first expressive end-to-end (E2E) spoken language model (SLM) that goes beyond natural conversation to support both role-playing and singing generation. VITA-QinYu adopts a hybrid speech-text paradigm that extends interleaved text-audio modeling with multi-codebook audio tokens, a design enabling richer paralinguistic representation while preserving a clear separation between modalities to avoid interference. We further develop a comprehensive data generation pipeline to synthesize a total of 15.8K hours of natural conversation, role-playing, and singing data for training. VITA-QinYu demonstrates superior expressiveness, outperforming peer SLMs by 7 percentage points on objective role-playing benchmarks, and surpassing peer models by 0.13 points on a 5-point MOS scale for singing. Simultaneously, it achieves state-of-the-art conversational accuracy and fluency, exceeding prior SLMs by 1.38 and 4.98 percentage points on the C3 and URO benchmarks, respectively. We open-source our code and models and provide an easy-to-use demo with full-stack support for streaming and full-duplex interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。