arXiv:2509.03883cs.CVcs.MM2025-09TPAMI综述被引 47

系统梳理人体动作视频生成全流程,涵盖视觉、文本、音频三模态技术。

Human Motion Video Generation: A Survey

  • 按输入到输出的五个阶段系统分类生成流程
  • 综述200+论文,覆盖视觉、文本、音频三类模态生成方法
  • 首次探讨大语言模型在动作生成中的潜力,适合数字人研究者

人体动作视频生成因广泛应用而备受关注,可实现逼真唱歌头像或随音乐起舞的动态虚拟形象。现有综述多聚焦单一方法,缺乏对完整生成流程的全面梳理。本文首次系统性地总结该领域,涵盖十多个子任务及五个关键阶段:输入、动作规划、视频生成、优化与输出。特别指出大语言模型在提升生成质量方面的潜在作用。综述覆盖视觉、文本、音频三大模态的最新进展,基于200余篇论文,梳理关键技术演进与里程碑工作,旨在揭示数字人应用前景,为后续研究提供参考。相关模型列表详见开源仓库:https://github.com/Winn1y/Awesome-Human-Motion-Video-Generation。

原文摘要 · Abstract (English)

Human motion video generation has garnered significant research interest due to its broad applications, enabling innovations such as photorealistic singing heads or dynamic avatars that seamlessly dance to music. However, existing surveys in this field focus on individual methods, lacking a comprehensive overview of the entire generative process. This paper addresses this gap by providing an in-depth survey of human motion video generation, encompassing over ten sub-tasks, and detailing the five key phases of the generation process: input, motion planning, motion video generation, refinement, and output. Notably, this is the first survey that discusses the potential of large language models in enhancing human motion video generation. Our survey reviews the latest developments and technological trends in human motion video generation across three primary modalities: vision, text, and audio. By covering over two hundred papers, we offer a thorough overview of the field and highlight milestone works that have driven significant technological breakthroughs. Our goal for this survey is to unveil the prospects of human motion video generation and serve as a valuable resource for advancing the comprehensive applications of digital humans. A complete list of the models examined in this survey is available in Our Repository https://github.com/Winn1y/Awesome-Human-Motion-Video-Generation.

动作生成数字人多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。