arXiv:2509.09595cs.CV2025-09被引 23

让虚拟人物动画更懂指令意图,生成更自然长视频。

Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis

  • 用多模态大模型理解指令,生成带情绪和动作的蓝图视频
  • 分阶段生成,支持1080p/48fps流畅长视频,口型同步精准
  • 适合数字人直播、Vlog等真实场景,可跨领域通用

近期音频驱动的虚拟人物视频生成技术显著提升了视听真实感。然而,现有方法仅将指令条件视为由声学或视觉线索驱动的低级追踪,未建模指令所传递的交流意图,导致叙事连贯性与角色表现力不足。为此,我们提出Kling-Avatar,一种统一多模态指令理解与逼真肖像生成的级联框架。该方法采用两阶段流程:第一阶段设计多模态大语言模型(MLLM)导演,基于多样化指令信号生成受控于高阶语义(如动作、情感)的蓝图视频;第二阶段在蓝图关键帧引导下,采用首尾帧策略并行生成多个子片段。此全局到局部框架在保留细粒度细节的同时,忠实编码多模态指令背后的高层意图。并行架构还实现长时视频的快速稳定生成,适用于数字人直播、Vlog等实际应用。为全面评估,我们构建包含375个精选样本的基准数据集,覆盖多样指令与挑战场景。大量实验表明,Kling-Avatar可生成高达1080p、48fps的生动流畅长视频,在口型同步精度、情绪动态表现力、指令可控性、身份保真度及跨域泛化能力上均表现卓越,确立其作为语义化、高保真音频驱动虚拟人合成的新标杆。

原文摘要 · Abstract (English)

Recent advances in audio-driven avatar video generation have significantly enhanced audio-visual realism. However, existing methods treat instruction conditioning merely as low-level tracking driven by acoustic or visual cues, without modeling the communicative purpose conveyed by the instructions. This limitation compromises their narrative coherence and character expressiveness. To bridge this gap, we introduce Kling-Avatar, a novel cascaded framework that unifies multimodal instruction understanding with photorealistic portrait generation. Our approach adopts a two-stage pipeline. In the first stage, we design a multimodal large language model (MLLM) director that produces a blueprint video conditioned on diverse instruction signals, thereby governing high-level semantics such as character motion and emotions. In the second stage, guided by blueprint keyframes, we generate multiple sub-clips in parallel using a first-last frame strategy. This global-to-local framework preserves fine-grained details while faithfully encoding the high-level intent behind multimodal instructions. Our parallel architecture also enables fast and stable generation of long-duration videos, making it suitable for real-world applications such as digital human livestreaming and vlogging. To comprehensively evaluate our method, we construct a benchmark of 375 curated samples covering diverse instructions and challenging scenarios. Extensive experiments demonstrate that Kling-Avatar is capable of generating vivid, fluent, long-duration videos at up to 1080p and 48 fps, achieving superior performance in lip synchronization accuracy, emotion and dynamic expressiveness, instruction controllability, identity preservation, and cross-domain generalization. These results establish Kling-Avatar as a new benchmark for semantically grounded, high-fidelity audio-driven avatar synthesis.

虚拟人生成多模态理解长视频合成指令控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。