arXiv:2602.07106cs.CVcs.AI2026-02被引 1

让大模型同时生成语音和3D人脸动画,提升人机交互自然度。

Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models

  • 用语音单元搭建时间骨架,分离语义与动作生成
  • 120万数据训练,音视频同步更准、延迟更低
  • 适合做虚拟助手、数字人等实时交互场景

全能模态大语言模型(OLLMs)旨在统一多模态理解与生成,但将其扩展至同步生成语音与3D人脸动画仍鲜有研究,而这对自然人机交互至关重要。核心挑战在于大模型的离散语义推理与3D人脸运动的密集时序动态之间的不匹配。本文提出Ex-Omni,一个开源模型,为OLLMs原生添加伴随语音的3D人脸动画生成能力。Ex-Omni通过一种关注混合形状的语音单元生成器与混合形状解码器,将语义推理与时序生成解耦:语音单元提供时序框架,隐藏语音表征携带面部相关线索。此外,引入统一的令牌即查询门控融合机制(TQGF),实现可控语义注入,并构建了包含1200K样本的InstructS2SF-1200K数据集用于预训练。大量实验表明,Ex-Omni在保持良好语音理解与生成能力的同时,实现了更优的音视频同步效果与更低的人脸生成延迟,优于级联式流水线方法。

原文摘要 · Abstract (English)

Omni-modal large language models (OLLMs) aim to unify multimodal understanding and generation, yet extending them to jointly produce speech and 3D facial animation remains largely unexplored despite its importance for natural human-computer interaction. A key challenge is the mismatch between the discrete semantic reasoning of LLMs and the dense temporal dynamics required for 3D facial motion. We propose Expressive Omni (Ex-Omni), an open-source model that augments OLLMs with native speech-accompanied 3D facial animation. Ex-Omni decouples semantic reasoning from temporal generation through a blendshape-aware speech unit generator and a blendshape decoder, where speech units provide temporal scaffolding and hidden speech representations carry facially relevant cues. We further introduce a unified token-as-query gated fusion (TQGF) mechanism for controlled semantic injection, as well as InstructS2SF-1200K, a dataset consisting of 1200K samples for pre-training. Extensive experiments show that Ex-Omni maintains competitive speech understanding and generation ability while achieving better audio-visual synchronization and lower face-generation latency than cascaded pipelines.

3D动画大模型语音生成数字人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。