UniTalker让对话语音与面部动画同步生成,更像真人互动。
UniTalker: Conversational Speech-Visual Synthesis
- 用大模型理解文本、语音和人脸动画多模态信息
- 生成情绪一致的语音和自然的说话人脸动画
- 适合虚拟助手、元宇宙交互等需要情感表达的场景
对话式语音合成(CSS)是人机交互中的关键任务,旨在生成更具表现力和共情能力的语音。然而,真实人际交流中“倾听”与“眼神接触”对情感传递至关重要。现有研究仅依赖对话中的文本与语音,限制了交互效果。为此,我们提出对话式音视频合成(CSVS)新任务,扩展传统CSS。我们构建了统一模型UniTalker,融合多模态感知与渲染能力:利用大语言模型理解说话人、文本、语音及说话人脸动画;通过多任务序列预测,先推断情绪,再生成共情语音与自然人脸动画。为保证音视频在情绪、内容与时长上一致,引入三项优化:1)设计专用神经地标编码器,实现面部表情序列的分词与重建;2)提出双模态语音-视觉硬对齐解码策略;3)生成阶段采用情绪引导渲染。客观与主观实验表明,该模型生成更具共情性的语音,并呈现更自然、情绪一致的说话人脸动画。
原文摘要 · Abstract (English)
Conversational Speech Synthesis (CSS) is a key task in the user-agent interaction area, aiming to generate more expressive and empathetic speech for users. However, it is well-known that "listening" and "eye contact" play crucial roles in conveying emotions during real-world interpersonal communication. Existing CSS research is limited to perceiving only text and speech within the dialogue context, which restricts its effectiveness. Moreover, speech-only responses further constrain the interactive experience. To address these limitations, we introduce a Conversational Speech-Visual Synthesis (CSVS) task as an extension of traditional CSS. By leveraging multimodal dialogue context, it provides users with coherent audiovisual responses. To this end, we develop a CSVS system named UniTalker, which is a unified model that seamlessly integrates multimodal perception and multimodal rendering capabilities. Specifically, it leverages a large-scale language model to comprehensively understand multimodal cues in the dialogue context, including speaker, text, speech, and the talking-face animations. After that, it employs multi-task sequence prediction to first infer the target utterance's emotion and then generate empathetic speech and natural talking-face animations. To ensure that the generated speech-visual content remains consistent in terms of emotion, content, and duration, we introduce three key optimizations: 1) Designing a specialized neural landmark codec to tokenize and reconstruct facial expression sequences. 2) Proposing a bimodal speech-visual hard alignment decoding strategy. 3) Applying emotion-guided rendering during the generation stage. Comprehensive objective and subjective experiments demonstrate that our model synthesizes more empathetic speech and provides users with more natural and emotionally consistent talking-face animations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。