arXiv:2604.05961cs.CV2026-04

用人体结构噪声生成更自然的人体视频,让动作与服装细节更真实。

HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation

  • 基于人体拓扑的3D关节噪声采样,确保时空一致的运动噪声。
  • 联合学习外观与物理运动,还原衣物褶皱等动态细节。
  • 无需修改模型架构,可直接用于图像到视频生成,适合可控视频创作。

尽管近期人像视频生成取得显著进展,但生成式视频扩散模型仍难以忠实捕捉人体动作的动力学与物理特性。本文提出新框架HumANDiff,通过三项关键设计提升人体运动控制能力:1)采用人体结构相关的时空噪声采样,将传统随机高斯噪声替换为在统计人体模板密集表面流形上采样的3D关节噪声,继承身体拓扑先验以实现时空一致性;2)联合外观-运动学习,通过从关节噪声中同时预测像素外观和对应物理运动,提升视频生成保真度,如准确呈现随动作变化的衣纹;3)几何运动一致性学习,定义于关节噪声空间的新损失函数,在帧间强制物理运动一致性。HumANDiff通过微调现有扩散模型即可实现可扩展的可控人体视频生成,对模型架构无依赖,且推理时支持单框架图像到视频生成,具备原生运动控制能力。大量实验表明,该方法在多样服饰下生成运动一致、高保真的人体视频方面达到当前最优性能。

原文摘要 · Abstract (English)

Despite tremendous recent progress in human video generation, generative video diffusion models still struggle to capture the dynamics and physics of human motions faithfully. In this paper, we propose a new framework for human video generation, HumANDiff, which enhances the human motion control with three key designs: 1) Articulated motion-consistent noise sampling that correlates the spatiotemporal distribution of latent noise and replaces the unstructured random Gaussian noise with 3D articulated noise sampled on the dense surface manifold of a statistical human body template. It inherits body topology priors for spatially and temporally consistent noise sampling. 2) Joint appearance-motion learning that enhances the standard training objective of video diffusion models by jointly predicting pixel appearances and corresponding physical motions from the articulated noises. It enables high-fidelity human video synthesis, e.g., capturing motion-dependent clothing wrinkles. 3) Geometric motion consistency learning that enforces physical motion consistency across frames via a novel geometric motion consistency loss defined in the articulated noise space. HumANDiff enables scalable controllable human video generation by fine-tuning video diffusion models with articulated noise sampling. Consequently, our method is agnostic to diffusion model design, and requires no modifications to the model architecture. During inference, HumANDiff enables image-to-video generation within a single framework, achieving intrinsic motion control without requiring additional motion modules. Extensive experiments demonstrate that our method achieves state-of-the-art performance in rendering motion-consistent, high-fidelity humans with diverse clothing styles. Project page: https://taohuumd.github.io/projects/HumANDiff/

视频生成扩散模型人体动作可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。