arXiv:2602.06053cs.CL2026-02被引 45

让语音对话模型能自由切换角色和声音,支持真实客服场景

PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models

  • 用角色提示+语音样本混合控制,实现双工对话中角色与声音的灵活切换
  • 在多角色客服场景中,角色遵循度、语音相似度和响应自然度均优于现有模型
  • 基于开源大模型生成海量训练数据,适用于个性化对话系统开发

近期双工语音模型实现了低延迟、自然的语音到语音交互。然而,现有模型受限于固定角色和声音,难以支持结构化、角色驱动的真实应用与个性化互动。本文提出PersonaPlex,一种融合角色条件与文本提示、语音克隆与语音样本的双工对话语音模型。该模型在大规模合成数据集上训练,数据由开源大语言模型(LLM)与文本转语音(TTS)模型生成。为评估真实场景中的角色控制能力,我们扩展了Full-Duplex-Bench基准,涵盖多角色客户服务场景。实验表明,PersonaPlex在角色遵循度、语音相似度、延迟与自然度方面均超越当前最优双工语音模型及基于混合大语言模型的语音系统。

原文摘要 · Abstract (English)

Recent advances in duplex speech models have enabled natural, low-latency speech-to-speech interactions. However, existing models are restricted to a fixed role and voice, limiting their ability to support structured, role-driven real-world applications and personalized interactions. In this work, we introduce PersonaPlex, a duplex conversational speech model that incorporates hybrid system prompts, combining role conditioning with text prompts and voice cloning with speech samples. PersonaPlex is trained on a large-scale synthetic dataset of paired prompts and user-agent conversations, generated with open-source large language models (LLM) and text-to-speech (TTS) models. To evaluate role conditioning in real-world settings, we extend the Full-Duplex-Bench benchmark beyond a single assistant role to multi-role customer service scenarios. Experiments show that PersonaPlex achieves strong role-conditioned behavior, voice-conditioned speech, and natural conversational responsiveness, surpassing state-of-the-art duplex speech models and hybrid large language model-based speech systems in role adherence, speaker similarity, latency, and naturalness.

语音对话角色控制双工交互语音克隆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。