测试三种姿态迁移模型生成动作的逼真度,发现真实感不足且动作一致性差。
Can Pose Transfer Models Generate Realistic Human Motion?
- 在跨身份和动作的测试中评估姿态迁移效果
- 仅42.92%观众能正确识别目标动作,36.46%认为动作与参考一致
- 基于点云融合的ExAvatar表现优于扩散模型方法
近期的姿态迁移方法旨在生成时间连贯、完全可控的人体动作视频,将参考视频中的动作迁移到新身份上。本文通过在训练分布外的动作与身份组合上生成视频,并开展参与者研究,评估三种前沿方法——AnimateAnyone、MagicAnimate和ExAvatar的表现。在20种不同人体动作的受控环境中,参与者仅42.92%能正确识别目标动作,且仅36.46%认为生成动作与参考视频一致。结果显示,基于点云融合(splatting)的ExAvatar在动作一致性和视觉真实感方面优于基于扩散模型的AnimateAnyone和MagicAnimate。
原文摘要 · Abstract (English)
Recent pose-transfer methods aim to generate temporally consistent and fully controllable videos of human action where the motion from a reference video is reenacted by a new identity. We evaluate three state-of-the-art pose-transfer methods -- AnimateAnyone, MagicAnimate, and ExAvatar -- by generating videos with actions and identities outside the training distribution and conducting a participant study about the quality of these videos. In a controlled environment of 20 distinct human actions, we find that participants, presented with the pose-transferred videos, correctly identify the desired action only 42.92% of the time. Moreover, the participants find the actions in the generated videos consistent with the reference (source) videos only 36.46% of the time. These results vary by method: participants find the splatting-based ExAvatar more consistent and photorealistic than the diffusion-based AnimateAnyone and MagicAnimate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。