arXiv:2602.00702cs.CV2026-02被引 1

让虚拟人更听话:通过文本和音频协同控制生成自然动作

JoyStreamer: Unlocking Highly Expressive Avatars via Harmonized Text-Audio Conditioning

  • 用双教师训练法融合文本与音视频同步控制
  • 动态调节多模态条件强度,减少信号冲突
  • 支持复杂动作、多人对话和非人类角色扮演

现有视频虚拟人模型在说话、演讲和唱歌等场景表现良好,但在处理包含大范围全身动作、动态摄像机轨迹、背景切换或人物与物体交互的复杂指令时,对文本指令的对齐能力有限。为此,我们提出JoyStreamer框架,可生成长时间的高表达力虚拟人视频。其核心创新包括:首先,采用双教师增强训练算法,使模型在保留基础模型文本可控性的同时学习音视频同步;其次,在训练中根据去噪步数动态调节多模态条件(如音频和文本)的强度,缓解异构条件信号间的冲突。两项设计显著提升了模型生成自然、时序连贯的全身运动和动态镜头移动的能力,同时保持唇形同步和身份一致性。GSB评估显示,本模型优于Omnihuman-1.5和KlingAvatar 2.0等当前最优模型。此外,该方法可实现多人对话及非人类角色扮演等复杂应用。部分视频样本见https://joystreamer.github.io/。

原文摘要 · Abstract (English)

Existing video avatar models have demonstrated impressive capabilities in scenarios such as talking, public speaking, and singing. However, the majority of these methods exhibit limited alignment with respect to text instructions, particularly when the prompts involve complex elements including large full-body movement, dynamic camera trajectory, background transitions, or human-object interactions. To break out this limitation, we present JoyAvatar, a framework capable of generating long duration avatar videos, featuring two key technical innovations. Firstly, we introduce a twin-teacher enhanced training algorithm that enables the model to transfer inherent text-controllability from the foundation model while simultaneously learning audio-visual synchronization. Secondly, during training, we dynamically modulate the strength of multi-modal conditions (e.g., audio and text) based on the distinct denoising timestep, aiming to mitigate conflicts between the heterogeneous conditioning signals. These two key designs serve to substantially expand the avatar model's capacity to generate natural, temporally coherent full-body motions and dynamic camera movements as well as preserve the basic avatar capabilities, such as accurate lip-sync and identity consistency. GSB evaluation results demonstrate that our JoyStreamer model outperforms the state-of-the-art models such as Omnihuman-1.5 and KlingAvatar 2.0. Moreover, our approach enables complex applications including multi-person dialogues and non-human subjects role-playing. Some video samples are provided on https://joystreamer.github.io/.

虚拟人多模态生成文本控制动作合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。