arXiv:2412.17290cs.CV2024-12CVPR被引 5

通过关联姿态的参考选择,实现人物动画的自由视角生成。

Free-viewpoint Human Animation with Pose-correlated Reference Selection

  • 根据姿态相似性动态选择最相关的参考图像
  • 在大幅视角变化下仍能保持高保真动画效果
  • 适合影视镜头规划与相机控制等需要自由视角的应用

基于扩散模型的人体动画能够根据源图像和驱动信号(如姿态序列)生成人物动作,但面对显著视角变化,尤其在推拉镜头(相机与人物距离变化)时,常因身体外观细节丢失而表现不佳,限制了其在电影镜头设计和相机控制中的应用。为此,本文提出一种姿态相关参考选择扩散网络,支持大幅度视角变化下的高质量人体动画生成。核心思路是允许多参考图像输入,以弥补视角变化带来的外观缺失。为降低计算开销,设计了一种新型姿态相关性模块,用于计算非对齐姿态间的相似性,并提出自适应参考选择策略,利用注意力图识别关键生成区域。为训练模型,我们从公开的TED演讲视频中构建了一个大规模数据集,包含同一人物的多种镜头视角。实验表明,在相同参考图像数量下,本方法在大视角变化场景下优于当前最优方法;且自适应选择机制能有效定位最相关的参考区域,生成自由视角下的人体图像。

原文摘要 · Abstract (English)

Diffusion-based human animation aims to animate a human character based on a source human image as well as driving signals such as a sequence of poses. Leveraging the generative capacity of diffusion model, existing approaches are able to generate high-fidelity poses, but struggle with significant viewpoint changes, especially in zoom-in/zoom-out scenarios where camera-character distance varies. This limits the applications such as cinematic shot type plan or camera control. We propose a pose-correlated reference selection diffusion network, supporting substantial viewpoint variations in human animation. Our key idea is to enable the network to utilize multiple reference images as input, since significant viewpoint changes often lead to missing appearance details on the human body. To eliminate the computational cost, we first introduce a novel pose correlation module to compute similarities between non-aligned target and source poses, and then propose an adaptive reference selection strategy, utilizing the attention map to identify key regions for animation generation. To train our model, we curated a large dataset from public TED talks featuring varied shots of the same character, helping the model learn synthesis for different perspectives. Our experimental results show that with the same number of reference images, our model performs favorably compared to the current SOTA methods under large viewpoint change. We further show that the adaptive reference selection is able to choose the most relevant reference regions to generate humans under free viewpoints.

人体动画扩散模型自由视角参考选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。