arXiv:2409.06791cs.CVcs.HC2024-09被引 2

用扩散模型自动补全人形动作,生成自然流畅的5秒运动序列。

Human Motion Synthesis_ A Diffusion Approach for Motion Stitching and In-Betweening

  • 基于Transformer的去噪器设计,实现动作片段间的平滑衔接。
  • 可处理任意数量输入姿态,生成75帧(5秒,15fps)真实感动作序列。
  • 适合动画、游戏等领域需要自动化动作补全的场景。

人体动作生成是多个领域的研究重点。本文针对动作拼接与中间帧生成问题,提出一种基于Transformer的去噪器扩散模型。现有方法要么依赖人工干预,要么无法处理长序列。所提方法能将任意数量的输入姿态转换为平滑且真实的动作序列,共75帧,帧率为15 fps,总时长5秒。通过弗雷切特起始距离(FID)、多样性与多模态性等量化指标,以及生成结果的视觉评估,验证了该方法在生成中间帧序列上的优异表现。

原文摘要 · Abstract (English)

Human motion generation is an important area of research in many fields. In this work, we tackle the problem of motion stitching and in-betweening. Current methods either require manual efforts, or are incapable of handling longer sequences. To address these challenges, we propose a diffusion model with a transformer-based denoiser to generate realistic human motion. Our method demonstrated strong performance in generating in-betweening sequences, transforming a variable number of input poses into smooth and realistic motion sequences consisting of 75 frames at 15 fps, resulting in a total duration of 5 seconds. We present the performance evaluation of our method using quantitative metrics such as Frechet Inception Distance (FID), Diversity, and Multimodality, along with visual assessments of the generated outputs.

动作生成扩散模型动作补全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。