arXiv:2510.17181cs.CV2025-10ICCV被引 1

从单目视频中还原带手部接触的逼真头部虚拟形象

Capturing Head Avatar with Hand Contacts from a Monocular Video

  • 联合学习面部与手部交互的非刚性形变,通过深度顺序和接触正则化保持空间关系
  • 在真实手机拍摄视频上实现更准确的面部几何重建,相比现有方法减少穿插伪影
  • 适用于虚拟会议、游戏等需自然手势交互的场景,对细节敏感度高

逼真3D头部虚拟形象在远程通讯、游戏和虚拟现实中有重要应用。然而,多数方法仅关注面部区域,忽略了手部与面部的自然交互(如手托下巴或手指轻触脸颊),这些动作能传达沉思等认知状态。本文提出一种新框架,联合学习高质量头部虚拟形象及手部接触引起的非刚性形变。主要挑战有两个:一是单独追踪手与脸无法捕捉其相对姿态,为此我们引入深度顺序损失与接触正则化以确保空间关系正确;二是缺乏公开的手部诱导形变先验,难以从单目视频中学习,因此我们从一个手脸交互数据集学习专门的主成分分析(PCA)基底,将问题简化为估计少量参数而非完整形变场。此外,借鉴物理模拟思想,引入接触损失提供额外监督,显著减少穿插伪影并提升结果的物理合理性。我们在iPhone采集的RGB(D)视频上评估该方法,并构建了包含多种手部交互类型的合成数据集用于几何评估。实验表明,本方法在外观和面部形变几何精度上优于当前最优表面重建方法。

原文摘要 · Abstract (English)

Photorealistic 3D head avatars are vital for telepresence, gaming, and VR. However, most methods focus solely on facial regions, ignoring natural hand-face interactions, such as a hand resting on the chin or fingers gently touching the cheek, which convey cognitive states like pondering. In this work, we present a novel framework that jointly learns detailed head avatars and the non-rigid deformations induced by hand-face interactions. There are two principal challenges in this task. First, naively tracking hand and face separately fails to capture their relative poses. To overcome this, we propose to combine depth order loss with contact regularization during pose tracking, ensuring correct spatial relationships between the face and hand. Second, no publicly available priors exist for hand-induced deformations, making them non-trivial to learn from monocular videos. To address this, we learn a PCA basis specific to hand-induced facial deformations from a face-hand interaction dataset. This reduces the problem to estimating a compact set of PCA parameters rather than a full spatial deformation field. Furthermore, inspired by physics-based simulation, we incorporate a contact loss that provides additional supervision, significantly reducing interpenetration artifacts and enhancing the physical plausibility of the results. We evaluate our approach on RGB(D) videos captured by an iPhone. Additionally, to better evaluate the reconstructed geometry, we construct a synthetic dataset of avatars with various types of hand interactions. We show that our method can capture better appearance and more accurate deforming geometry of the face than SOTA surface reconstruction methods.

三维重建手部交互单目视频虚拟形象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。