arXiv:2510.03909cs.CV2025-10被引 5

用文本生成真人动作视频,跨阶段对齐更精准。

Generating Human Motion Videos using a Cascaded Text-to-Video Framework

  • 分阶段融合文本到动作与视频生成模型,优化对齐效果。
  • 引入相机感知模块,自动选视角提升画面连贯性。
  • 支持电影级视频生成,适用场景广泛。

人类视频生成在图形学、娱乐和具身AI中有广泛应用。尽管视频扩散模型(VDMs)进展迅速,但其在通用真人动作视频生成中的应用仍不充分,多数研究局限于图像到视频或舞蹈等特定领域。本文提出CAMEO,一种用于通用人类动作视频生成的级联框架。该框架无缝连接文本到动作(T2M)模型与条件视频扩散模型(VDM),通过精心设计的组件缓解训练与推理中可能出现的性能劣化问题。具体而言,我们分析并优化了文本提示与视觉条件,以有效训练VDM,确保动作描述、条件信号与生成视频之间的稳健对齐。此外,引入相机感知条件模块,在两阶段间建立联系,自动选择与输入文本匹配的视角,增强一致性并减少人工干预。我们在MovieGen基准及一个新提出的专为T2M-VDM组合设计的基准上验证了方法的有效性,并展示了其在多种应用场景下的多样性与实用性。

原文摘要 · Abstract (English)

Human video generation is becoming an increasingly important task with broad applications in graphics, entertainment, and embodied AI. Despite the rapid progress of video diffusion models (VDMs), their use for general-purpose human video generation remains underexplored, with most works constrained to image-to-video setups or narrow domains like dance videos. In this work, we propose CAMEO, a cascaded framework for general human motion video generation. It seamlessly bridges Text-to-Motion (T2M) models and conditional VDMs, mitigating suboptimal factors that may arise in this process across both training and inference through carefully designed components. Specifically, we analyze and prepare both textual prompts and visual conditions to effectively train the VDM, ensuring robust alignment between motion descriptions, conditioning signals, and the generated videos. Furthermore, we introduce a camera-aware conditioning module that connects the two stages, automatically selecting viewpoints aligned with the input text to enhance coherence and reduce manual intervention. We demonstrate the effectiveness of our approach on both the MovieGen benchmark and a newly introduced benchmark tailored to the T2M-VDM combination, while highlighting its versatility across diverse use cases.

视频生成文本到视频动作生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。