arXiv:2604.17211cs.CV2026-04被引 1

让大模型对话时有实时逼真口型动作,支持无缝换人说话。

EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational Agents

论文配图:EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational Agents
图 1 · 摘自论文原文
  • 用单流音频+状态控制,实现听和说同步,不提前看对方话
  • 4步生成高质量口型,支持实时交互,口型动作自然无误
  • 适合做虚拟助手、数字人,尤其需要真实对话感的场景

我们提出EmbodiedHead,一个语音驱动的实时对话头像框架,为大语言模型赋予视觉化身。真实可用的具身化身需同时满足实时生成、听-说行为统一和高画质渲染。该框架首次采用修正流扩散Transformer(DiT)结合可微分渲染器,仅需4步采样即可生成多样且高保真的口型动作。以往方法依赖双流音频,导致听者提前预判发言者内容,与因果对话机制冲突。本文改用单流接口,通过显式帧级听-说状态条件和流式音频调度器,在倾听时抑制虚假嘴动,实现无缝换人对话。采用两阶段训练策略:系数空间预训练与图像域联合优化,有效弥合运动级监督与渲染质量之间的差距。大量实验表明,该方法在发声与倾听场景下均达到当前最佳视觉质量和动作一致性。

原文摘要 · Abstract (English)

We present EmbodiedHead, a speech-driven talking-head framework that equips LLMs with real-time visual avatars for conversation. A practical embodied avatar must achieve real-time generation, unified listening-speaking behavior, and high rendered visual quality simultaneously. Our framework couples the first Rectified-Flow Diffusion Transformer (DiT) for this task with a differentiable renderer, enabling diverse, high-fidelity generation in as few as four sampling steps. Prior listening-speaking methods rely on dual-stream audio, introducing an interlocutor look-ahead dependency incompatible with causal user--LLM interaction. We instead adopt a single-stream interface with explicit per-frame listening-speaking state conditioning and a Streaming Audio Scheduler, suppressing spurious mouth motion during listening while enabling seamless turn-taking. A two-stage training scheme of coefficient-space pretraining and joint image-domain refinement further closes the gap between motion-level supervision and rendered quality. Extensive experiments demonstrate state-of-the-art visual quality and motion fidelity in both speaking and listening scenarios.

数字人语音驱动实时生成对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。