arXiv:2505.20277cs.CLcs.CV2025-05ACL被引 20

让角色扮演智能体同时具备语音与语言的个性表现,实现沉浸式交互。

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction

  • 统一建模角色语言与语音个性,支持语音与文字混合响应。
  • 在10K对话数据上实现289ms低延迟,内容与风格均优于现有模型。
  • 适合需要高度拟人化交互的场景,如虚拟陪伴、游戏NPC。

角色扮演智能体(RPAs)借助大语言模型成为新兴的交互式AI系统,可模拟具有多样人格的角色。然而现有方法多聚焦于文本对话模仿,忽视了语音特征(如音色、情感)在真实交互中对沉浸感的关键作用。为此,我们提出OmniCharacter——首个实现无缝语音-语言人格交互的模型,支持角色在交互中持续展现特定人格与语音特质,并生成语音与文字混合响应。为适配该任务,我们构建了OmniCharacter-10K数据集,包含20个独特角色、10,000轮丰富语境对话及135,000条动态语音响应。实验表明,该方法在内容与风格上均优于现有RPAs及主流语音-语言模型,响应延迟低至289ms。代码与数据已开源。

原文摘要 · Abstract (English)

Role-Playing Agents (RPAs), benefiting from large language models, is an emerging interactive AI system that simulates roles or characters with diverse personalities. However, existing methods primarily focus on mimicking dialogues among roles in textual form, neglecting the role's voice traits (e.g., voice style and emotions) as playing a crucial effect in interaction, which tends to be more immersive experiences in realistic scenarios. Towards this goal, we propose OmniCharacter, a first seamless speech-language personality interaction model to achieve immersive RPAs with low latency. Specifically, OmniCharacter enables agents to consistently exhibit role-specific personality traits and vocal traits throughout the interaction, enabling a mixture of speech and language responses. To align the model with speech-language scenarios, we construct a dataset named OmniCharacter-10K, which involves more distinctive characters (20), richly contextualized multi-round dialogue (10K), and dynamic speech response (135K). Experimental results showcase that our method yields better responses in terms of both content and style compared to existing RPAs and mainstream speech-language models, with a response latency as low as 289ms. Code and dataset are available at https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/OmniCharacter.

角色扮演语音生成多模态交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。