arXiv:2505.22046cs.CV2025-05被引 2

针对复杂人体动作生成难题,提出新框架提升视频真实感。

LatentMove: Towards Complex Human Movement Video Generation

  • 基于DiT架构,引入条件控制分支与可学习体面标记
  • 在快速复杂动作上显著减少变形,提升运动一致性
  • 适合研究视频生成、动作模拟的开发者与研究人员

图像到视频(I2V)生成旨在从单张参考图像生成逼真的运动序列。尽管近期方法在时间一致性方面表现良好,但在处理复杂、非重复的人体动作时仍易出现不自然形变。为此,我们提出LatentMove,一种专为高动态人体动画设计的DiT-based框架。其架构包含条件控制分支和可学习的脸部/身体标记,以保持帧间一致性和细节精度。我们构建了Complex-Human-Videos(CHV)数据集,包含多样化、具挑战性的人体动作,用于评估I2V系统的鲁棒性。同时提出两项新指标,用于衡量生成视频与真实视频在运动流与轮廓一致性上的差异。实验表明,LatentMove在处理快速、复杂动作时显著提升了动画质量,推动了I2V生成的边界。代码、CHV数据集及评估指标将公开于https://github.com/--。

原文摘要 · Abstract (English)

Image-to-video (I2V) generation seeks to produce realistic motion sequences from a single reference image. Although recent methods exhibit strong temporal consistency, they often struggle when dealing with complex, non-repetitive human movements, leading to unnatural deformations. To tackle this issue, we present LatentMove, a DiT-based framework specifically tailored for highly dynamic human animation. Our architecture incorporates a conditional control branch and learnable face/body tokens to preserve consistency as well as fine-grained details across frames. We introduce Complex-Human-Videos (CHV), a dataset featuring diverse, challenging human motions designed to benchmark the robustness of I2V systems. We also introduce two metrics to assess the flow and silhouette consistency of generated videos with their ground truth. Experimental results indicate that LatentMove substantially improves human animation quality--particularly when handling rapid, intricate movements--thereby pushing the boundaries of I2V generation. The code, the CHV dataset, and the evaluation metrics will be available at https://github.com/ --.

视频生成人体动作扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。