无需标记和标定,实现全身高精度动作捕捉。
Look Ma, no markers: holistic performance capture without the hassle
- 结合合成数据训练的机器学习与人体参数模型,实现端到端重建。
- 在多种相机布置和服装下稳定输出世界坐标系结果。
- 首次支持面部、身体、手部及眼球、舌头的同步高精度捕捉。
本文解决面部、身体和手部同时进行高精度全身心动捕捉的问题。现有影视与游戏制作中的动捕技术通常只针对某一部分,依赖复杂昂贵的硬件和大量人工操作。现有基于机器学习的方法多仅支持单摄像头,仅处理身体局部,无法获得精确的世界空间结果,且泛化能力差。本工作提出首个无需标记、无需标定、无需定制硬件的全身心动捕捉方法,可从任意相机阵列稳定生成世界空间结果,适应不同环境与着装。通过结合仅在合成数据上训练的机器学习模型与强大的人体形状与运动参数模型,实现了对眼睛、舌头等细节的高精度重建。我们在多个面部、身体和手部重建基准上评估,结果表明该方法在多样化数据集上均达到领先性能并具备良好泛化能力。
原文摘要 · Abstract (English)
We tackle the problem of highly-accurate, holistic performance capture for the face, body and hands simultaneously. Motion-capture technologies used in film and game production typically focus only on face, body or hand capture independently, involve complex and expensive hardware and a high degree of manual intervention from skilled operators. While machine-learning-based approaches exist to overcome these problems, they usually only support a single camera, often operate on a single part of the body, do not produce precise world-space results, and rarely generalize outside specific contexts. In this work, we introduce the first technique for marker-free, high-quality reconstruction of the complete human body, including eyes and tongue, without requiring any calibration, manual intervention or custom hardware. Our approach produces stable world-space results from arbitrary camera rigs as well as supporting varied capture environments and clothing. We achieve this through a hybrid approach that leverages machine learning models trained exclusively on synthetic data and powerful parametric models of human shape and motion. We evaluate our method on a number of body, face and hand reconstruction benchmarks and demonstrate state-of-the-art results that generalize on diverse datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。