arXiv:2602.21188cs.CV2026-02被引 1

用单图生成可控视角与动作的高质量人体视频

Human Video Generation from a Single Image with 3D Pose and View Control

  • 通过3D骨骼图捕捉人体结构,解决跨视角自遮挡问题
  • 多视角一致且帧间稳定,生成长时序连贯动画
  • 适合需要精细动作与视角控制的人体视频生成场景

近期扩散模型在图像到视频生成方面取得显著进展,但人体视频生成仍面临挑战,尤其难以从单张图像推断出与运动相关的视图一致的衣物褶皱。本文提出4D人体视频生成模型HVG,可在3D姿态和视角控制下,从单张图像生成高质量、多视角、时空一致的人体视频。HVG通过三项关键设计实现:(i) 装备新型双维骨骼图的关节调制,利用3D信息捕捉人体解剖关系并解决跨视角自遮挡;(ii) 视角与时间对齐机制,确保多视角一致性及参考图像与姿态序列间的帧间稳定性;(iii) 分阶段时空采样结合时间对齐,保持长序列多视角动画的平滑过渡。大量实验表明,HVG在多种人体图像与姿态输入下均优于现有方法,可生成高质量4D人体视频。

原文摘要 · Abstract (English)

Recent diffusion methods have made significant progress in generating videos from single images due to their powerful visual generation capabilities. However, challenges persist in image-to-video synthesis, particularly in human video generation, where inferring view-consistent, motion-dependent clothing wrinkles from a single image remains a formidable problem. In this paper, we present Human Video Generation in 4D (HVG), a latent video diffusion model capable of generating high-quality, multi-view, spatiotemporally coherent human videos from a single image with 3D pose and view control. HVG achieves this through three key designs: (i) Articulated Pose Modulation, which captures the anatomical relationships of 3D joints via a novel dual-dimensional bone map and resolves self-occlusions across views by introducing 3D information; (ii) View and Temporal Alignment, which ensures multi-view consistency and alignment between a reference image and pose sequences for frame-to-frame stability; and (iii) Progressive Spatio-Temporal Sampling with temporal alignment to maintain smooth transitions in long multi-view animations. Extensive experiments on image-to-video tasks demonstrate that HVG outperforms existing methods in generating high-quality 4D human videos from diverse human images and pose inputs.

视频生成扩散模型3D姿态人体建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。