用3D人体结构引导扩散模型,实现可控的人体关键帧插值。
Controllable Human-centric Keyframe Interpolation with Generative Prior
- 引入3D人体姿态作为条件信号,融合到扩散过程中。
- 在新数据集上提升PSNR 9%,降低LPIPS 38%。
- 适合需要精确人体动作控制的视频生成任务。
现有插值方法依赖预训练视频扩散先验,在稀疏关键帧间生成中间帧。缺乏3D几何引导时,难以处理复杂人体动作,且对合成动态控制有限。本文提出PoseFuse3D关键帧插值器(PoseFuse3D-KI),将3D人体引导信号融入扩散过程,实现可控人体中心插值(CHKI)。为提供丰富空间与结构线索,提出的PoseFuse3D采用新型SMPL-X编码器,将3D几何与形状映射至2D潜在条件空间,并通过融合网络整合3D线索与2D姿态嵌入。我们构建了包含2D姿态与3D SMPL-X参数标注的新数据集CHKI-Video。实验表明,PoseFuse3D-KI在该数据集上持续优于当前最优基线,PSNR提升9%,LPIPS下降38%。全面消融实验证明,所提PoseFuse3D显著提升插值保真度。
原文摘要 · Abstract (English)
Existing interpolation methods use pre-trained video diffusion priors to generate intermediate frames between sparsely sampled keyframes. In the absence of 3D geometric guidance, these methods struggle to produce plausible results for complex, articulated human motions and offer limited control over the synthesized dynamics. In this paper, we introduce PoseFuse3D Keyframe Interpolator (PoseFuse3D-KI), a novel framework that integrates 3D human guidance signals into the diffusion process for Controllable Human-centric Keyframe Interpolation (CHKI). To provide rich spatial and structural cues for interpolation, our PoseFuse3D, a 3D-informed control model, features a novel SMPL-X encoder that transforms 3D geometry and shape into the 2D latent conditioning space, alongside a fusion network that integrates these 3D cues with 2D pose embeddings. For evaluation, we build CHKI-Video, a new dataset annotated with both 2D poses and 3D SMPL-X parameters. We show that PoseFuse3D-KI consistently outperforms state-of-the-art baselines on CHKI-Video, achieving a 9% improvement in PSNR and a 38% reduction in LPIPS. Comprehensive ablations demonstrate that our PoseFuse3D model improves interpolation fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。