arXiv:2412.14484cs.CV2024-12被引 8

用大模型当导演,让视频中人物动作更自然真实。

Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation

  • 让大模型生成人体姿态指令,指导视频生成
  • 在多个评测中人形动作真实度、贴合提示度均领先
  • 可适配不同视频渲染器,适合做角色驱动视频生成

本文提出DirectorLLM,一种新型视频生成模型,利用大语言模型(LLM)协调视频中的人体姿态。随着基础文本到视频模型快速发展,对高质量人体运动与交互的需求日益增长。为提升人体动作的真实性,我们将LLM从文本生成器扩展为视频导演和人体运动模拟器。基于开源Llama 3资源,训练DirectorLLM生成详细指令信号(如人体姿态),用于引导视频生成。该方法将人体运动模拟任务从视频生成器转移至LLM,生成具有信息量的人类中心场景草图。这些信号作为条件输入视频渲染器,实现更真实、更符合提示的视频生成。作为独立的LLM模块,可轻松适配多种渲染器(如UNet与DiT),实验显示其在自动评估与人工评估中均优于现有方法,在人体动作保真度、提示忠实度与主体自然度方面表现更佳。

原文摘要 · Abstract (English)

In this paper, we introduce DirectorLLM, a novel video generation model that employs a large language model (LLM) to orchestrate human poses within videos. As foundational text-to-video models rapidly evolve, the demand for high-quality human motion and interaction grows. To address this need and enhance the authenticity of human motions, we extend the LLM from a text generator to a video director and human motion simulator. Utilizing open-source resources from Llama 3, we train the DirectorLLM to generate detailed instructional signals, such as human poses, to guide video generation. This approach offloads the simulation of human motion from the video generator to the LLM, effectively creating informative outlines for human-centric scenes. These signals are used as conditions by the video renderer, facilitating more realistic and prompt-following video generation. As an independent LLM module, it can be applied to different video renderers, including UNet and DiT, with minimal effort. Experiments on automatic evaluation benchmarks and human evaluations show that our model outperforms existing ones in generating videos with higher human motion fidelity, improved prompt faithfulness, and enhanced rendered subject naturalness.

视频生成大模型动作模拟人类姿态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。