用四类智能体生成高保真、动作自然的人类动态视频。
HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics
- 分四步:重建3D人体场景、反思优化细节、生成表情化动作、合成逼真视频。
- 在文本生成、视频重演等任务上超越现有方法,几何一致性提升明显。
- 适合做虚拟人、影视特效或交互式内容生成的研究者和开发者。
合成人类动态旨在生成表现力强、意图驱动的逼真人像视频。当前方法面临两大挑战:(1) 几何不一致与粗略重建,源于有限的3D建模和细节保留;(2) 动作泛化能力弱与场景不协调,归因于生成能力不足。为此,我们提出HumanGenesis框架,通过四个协作智能体实现几何与生成建模融合:(1) 重建器利用单目视频与3D高斯点云分解,构建一致的3D人-场景表示;(2) 批评智能体通过多轮基于大语言模型的反思,识别并优化低质区域;(3) 姿态引导器使用时间感知参数编码器生成富有表现力的姿态序列;(4) 视频调和器采用混合渲染管线结合扩散模型,合成逼真连贯视频,并通过回环反馈优化重建器。HumanGenesis在文本引导生成、视频重演和新姿态泛化等任务上达到当前最优,显著提升表现力、几何保真度与场景融合性。
原文摘要 · Abstract (English)
\textbf{Synthetic human dynamics} aims to generate photorealistic videos of human subjects performing expressive, intention-driven motions. However, current approaches face two core challenges: (1) \emph{geometric inconsistency} and \emph{coarse reconstruction}, due to limited 3D modeling and detail preservation; and (2) \emph{motion generalization limitations} and \emph{scene inharmonization}, stemming from weak generative capabilities. To address these, we present \textbf{HumanGenesis}, a framework that integrates geometric and generative modeling through four collaborative agents: (1) \textbf{Reconstructor} builds 3D-consistent human-scene representations from monocular video using 3D Gaussian Splatting and deformation decomposition. (2) \textbf{Critique Agent} enhances reconstruction fidelity by identifying and refining poor regions via multi-round MLLM-based reflection. (3) \textbf{Pose Guider} enables motion generalization by generating expressive pose sequences using time-aware parametric encoders. (4) \textbf{Video Harmonizer} synthesizes photorealistic, coherent video via a hybrid rendering pipeline with diffusion, refining the Reconstructor through a Back-to-4D feedback loop. HumanGenesis achieves state-of-the-art performance on tasks including text-guided synthesis, video reenactment, and novel-pose generalization, significantly improving expressiveness, geometric fidelity, and scene integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。