arXiv:2604.19720cs.CV2026-04

用图像先验生成高质量可控人像视频,突破动作与视角限制。

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis

论文配图:ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis
图 1 · 摘自论文原文
  • 先生成高质量图像,再引导视频时序一致性
  • 支持多样姿态与视角,视频质量显著提升
  • 无需训练即可优化,适合影视与虚拟人开发

人像视频生成因需联合建模外观、运动与相机视角,在多视角数据有限下仍具挑战。现有方法常分步处理,导致控制力弱或画质下降。本文提出图像优先范式:先通过图像生成学习高质量外观,并作为视频合成先验,解耦外观建模与时序一致性。设计一个可控制姿态与视角的流程,结合预训练图像主干与基于SMPL-X的运动引导,外加基于预训练视频扩散模型的免训练时序优化阶段。实验表明,该方法在多种姿态与视角下生成高质量、时序一致的视频。同时发布一个标准人体数据集及用于组合式人像生成的辅助模型。代码与数据已公开于 https://github.com/Taited/ReImagine。

原文摘要 · Abstract (English)

Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint under limited multi-view data. Existing methods often address these factors separately, resulting in limited controllability or reduced visual quality. We revisit this problem from an image-first perspective, where high-quality human appearance is learned via image generation and used as a prior for video synthesis, decoupling appearance modeling from temporal consistency. We propose a pose- and viewpoint-controllable pipeline that combines a pretrained image backbone with SMPL-X-based motion guidance, together with a training-free temporal refinement stage based on a pretrained video diffusion model. Our method produces high-quality, temporally consistent videos under diverse poses and viewpoints. We also release a canonical human dataset and an auxiliary model for compositional human image synthesis. Code and data are publicly available at https://github.com/Taited/ReImagine.

人像视频图像先验扩散模型姿态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。