开源模型首次实现视觉语音文本三模态情感对话
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
- 用语义-声学解耦的语音分词器,统一处理多模态输入输出
- 在视觉语言和语音任务上均达到最新最优表现
- 支持情感丰富、可调音调的自然语音对话,适合语音交互应用
GPT-4o 作为具备多模态能力的通用模型,实现了带有情感与语调的语音对话,标志着多模态基础模型的重要进展。然而,对于开源社区而言,仅用公开数据实现大语言模型端到端感知与生成图像、文本及语音仍具挑战。现有视觉语言模型依赖外部工具进行语音处理,而语音语言模型则普遍缺乏视觉理解能力。为此,我们提出 EMOVA(情感全感知语音助手),使大语言模型具备端到端语音能力,同时保持领先的视觉语言性能。通过语义-声学解耦的语音分词器,我们发现多模态对齐能进一步提升视觉语言与语音能力,优于双模态对齐模型。此外,引入轻量级风格模块,实现情感与音调的灵活控制。EMOVA 首次在视觉语言与语音基准测试中均达到当前最佳表现,并支持充满情感的多模态语音对话。
原文摘要 · Abstract (English)
GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language Models to perceive and generate images, texts, and speeches end-to-end with publicly available data remains challenging for the open-source community. Existing vision-language models rely on external tools for speech processing, while speech-language models still suffer from limited or totally without vision-understanding capabilities. To address this gap, we propose the EMOVA (EMotionally Omni-present Voice Assistant), to enable Large Language Models with end-to-end speech abilities while maintaining the leading vision-language performance. With a semantic-acoustic disentangled speech tokenizer, we surprisingly notice that omni-modal alignment can further enhance vision-language and speech abilities compared with the bi-modal aligned counterparts. Moreover, a lightweight style module is introduced for the flexible speech style controls including emotions and pitches. For the first time, EMOVA achieves state-of-the-art performance on both the vision-language and speech benchmarks, and meanwhile, supporting omni-modal spoken dialogue with vivid emotions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。