arXiv:2608.01119cs.SD2026-08被引 1

让语音助手能懂情绪、会回应,实现自然流畅的双向对话。

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents

论文配图:JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents
图 1 · 摘自论文原文
  • 分模块设计思考-说话架构,联合训练保持文本推理能力
  • 支持笑声、叹气等细节控制,生成更富表现力的语音响应
  • 可感知用户情绪年龄等特征,适合需要共情的智能客服场景

我们提出 JoyAI-Talker,一个全双工语音对话系统,具备强大的基础模型能力,并支持共情交互与语音代理智能。该系统采用模块化 Thinker-Talker 架构,引入统一的语音-文本联合训练流程,缓解常见的“认知退化”瓶颈,有效保留模型在文本推理、STEM 和逻辑能力上的核心性能,并将其拓展至语音交互。在语音合成方面,Talker 模块采用文本可控生成范式,使自然语言指令可灵活控制音色、局部副语言事件(如笑声、叹气),支持更丰富精细的语音响应。为提升对话共情能力,引入层级式情感响应框架 PAER:从原始音频中提取性别、年龄、情绪状态等非言语线索,融入 Thinker 的思维链推理,生成语义适配且在话语级表达与局部副语言事件(如叹息、语速、音量)上具备精细控制的响应。进一步集成 Joy-Duplex 全双工框架,作为状态驱动的即插即用式实时换位控制引擎。大量评估显示,JoyAI-Talker 在基础的 T2T 与 S2T 基准上表现优异;在全双工测试中,用户打断下响应率达 0.88,背景语音下的误触发率极低,展现出流畅自然语音对话的实用潜力。

原文摘要 · Abstract (English)

We present JoyAI-Talker, a full-duplex speech dialogue system that delivers robust foundation model capabilities while empowering empathetic interaction and voice agent intelligence. JoyAI-Talker adopts a modular Thinker-Talker architecture and further implements a unified speech-text joint training pipeline to mitigate the common "cognitive degradation" bottleneck, thereby largely preserving the model's core textual reasoning, STEM, and logical capabilities while extending them to speech-based interaction. For expressive speech synthesis, the Talker module employs a text-controllable generation paradigm that enables natural-language instructions to flexibly control vocal attributes and localized paralinguistic events, such as laughter and sighs, supporting more expressive and fine-grained speech responses. To enhance conversational empathy, we introduce the Persona-Adaptive Empathetic Response (PAER) framework. PAER employs a hierarchical cognitive pipeline to extract non-verbal speaker cues, such as gender, age, and emotional state, from raw input audio, incorporate them into the Thinker's CoT reasoning, and generate context-adaptive responses that align semantically appropriate text with fine-grained control over utterance-level expressiveness and localized paralinguistic events, including sighs, speaking rate, and volume. We further integrate Joy-Duplex, a state-driven, plug-and-play full-duplex framework that functions as an efficient gating engine for real-time turn control. Extensive evaluations show that JoyAI-Talker achieves highly competitive performance on foundational T2T and S2T benchmarks. In full-duplex evaluation, the system reaches a high response rate of 0.88 under user interruptions while maintaining an extremely low false-trigger rate under background speech, demonstrating its readiness for fluid and natural speech dialogue.

语音交互共情对话全双工语音生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。