用约20个视角实现高保真手部动态捕捉,兼顾几何精度与自遮挡处理。
VEPHand: View-Efficient Photometric Hand Performance Capture at Scale

- 无掩码神经方法结合场景参数化与密度正则,提升有限视角下的手形与外观重建
- 基于物理启发的框架优化体积偏移与姿态,精准捕捉复杂形变与自接触
- 适用于单手、双手交互及自然抓握,支持大规模自动化采集
高保真3D手部捕捉对数字人创建至关重要,但实际多视角系统在有限视角密度下仍面临光照细节与几何歧义的挑战。本文提出端到端的手部动态性能捕捉与配准流水线,专为约20个视角的视点高效设置设计。通过两项核心创新克服关键难题:首先,采用无需掩码的神经方法,利用场景参数化和特定场景密度正则,在未遮挡图像中鲁棒提取精细手部几何与外观;其次,针对非线性皮肤变形与严重自遮挡下的配准问题,提出物理启发框架,通过优化个性化手模型的规则四面体网格内的体素偏移及姿态参数,实现高精度匹配。该方法结合鲁棒损失与优化策略,可准确捕捉微表面形变,确保极端动作下的合理性,并对输入噪声具有强容忍度。我们在超过12,000段序列上验证了流水线的可扩展性与鲁棒性,并由此构建了一个大规模高质量合成2D/3D手部数据集,用于下游任务训练。实验表明,该方法在视点高效、无掩码场景下达到当前最优重建保真度与注册精度,适用于单手、复杂双手交互及自然手物操作。
原文摘要 · Abstract (English)
Robust, high-fidelity 3D hand capture, while fundamental to digital human creation, remains challenging with practical multi-view systems that balance rich photometry with the geometric ambiguities of reconstruction arising from limited viewpoint density. This paper presents an end-to-end pipeline for dynamic hand performance capture and registration, specifically designed for view-efficient setups ($\sim$20 views). We address key challenges with two primary innovations. First, to overcome reconstruction difficulties like limited view overlap and background clutter, our mask-free neural method robustly extracts detailed hand geometry and appearance from unmasked images using scene parameterization and scenario-specific density regularization. Second, addressing registration challenges such as accurately capturing non-linear skin deformations and ensuring plausible results during severe self-contact, we propose a physics-inspired framework. It aligns reconstructions to a personalized hand model by optimizing intrinsic volumetric offsets within its canonical tetrahedral mesh, alongside pose parameters. This approach, supported by robust losses and optimization, captures fine surface deformations, ensures plausible results under severe articulation and self-contact, and demonstrates strong tolerance to input noise. We demonstrate the scalability and robustness of our automated pipeline on an extensive dataset of over 12,000 sequences, from which we also derive a large-scale, high-quality synthetic 2D/3D hand dataset for training downstream tasks. This showcases its effectiveness for single hands, intricate two-hand interactions, and natural hand-object manipulations. Our method achieves state-of-the-art reconstruction fidelity in view-efficient, unmasked scenarios and highly accurate registration. Our project page are available at https://vephand.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。