arXiv:2409.16666cs.CV2024-09ECCV被引 7

用单目视频生成带手部和面部动作的全身说话人动态3D模型

TalkinNeRF: Animatable Neural Fields for Full-Body Talking Humans

  • 统一神经辐射场建模全身4D运动,分模块学习身体、面部和手部姿态
  • 支持未见姿势下稳定动画,仅需短视频即可泛化到新身份
  • 可精细还原手指关节运动,适合影视与虚拟人制作场景

我们提出一种新框架,从单目视频中学习全身体态说话人的动态神经辐射场(NeRF)。以往工作仅建模身体姿态或面部表情,但人类交流依赖全身动作,包括肢体姿态、手势和面部表情。本文提出TalkinNeRF,一个统一的基于NeRF的网络,用于表示完整的4D人体运动。给定单个视频,我们学习身体、面部和手部对应的模块,并将其融合生成最终结果。为捕捉复杂的手指关节运动,额外学习手部形变场。多身份建模支持多个主体的同时训练,并在完全未见的姿态下保持鲁棒性。该方法可仅凭短视频输入泛化到新身份。实验表明,在生成全身说话人动画方面达到当前最佳性能,具备精细的手部关节运动与面部表情还原能力。

原文摘要 · Abstract (English)

We introduce a novel framework that learns a dynamic neural radiance field (NeRF) for full-body talking humans from monocular videos. Prior work represents only the body pose or the face. However, humans communicate with their full body, combining body pose, hand gestures, as well as facial expressions. In this work, we propose TalkinNeRF, a unified NeRF-based network that represents the holistic 4D human motion. Given a monocular video of a subject, we learn corresponding modules for the body, face, and hands, that are combined together to generate the final result. To capture complex finger articulation, we learn an additional deformation field for the hands. Our multi-identity representation enables simultaneous training for multiple subjects, as well as robust animation under completely unseen poses. It can also generalize to novel identities, given only a short video as input. We demonstrate state-of-the-art performance for animating full-body talking humans, with fine-grained hand articulation and facial expressions.

神经辐射场全身动画手势生成单目视频重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。