arXiv:2601.22693cs.CVcs.AI2026-01International Conf…被引 1

快速精准恢复单张照片中的人体网格,尤其擅长面部和手部细节。

PEAR: Pixel-aligned Expressive humAn mesh Recovery

  • 用统一视觉变换器模型实现高速推理,避免复杂架构
  • 引入像素级监督,显著提升面部与手部等细节重建精度
  • 无需预处理,每秒可生成100+帧,适合实时应用

从一张野外图像中重建精细的3D人体网格仍是计算机视觉中的基础挑战。现有基于SMPLX的方法常因推理缓慢、仅能生成粗略姿态,且在面部和手部等细微区域出现错位或不自然伪影,难以用于下游任务。为解决这些问题,我们提出PEAR——一种快速且鲁棒的像素对齐表达性人体网格恢复框架。PEAR明确应对三大缺陷:推理速度慢、细粒度姿态定位不准、面部表情捕捉不足。为实现实时SMPLX参数推断,我们摒弃依赖高分辨率输入或多分支结构的旧设计,采用简洁统一的ViT模型,可恢复粗略3D人体几何。为弥补此简化架构导致的细节损失,我们引入像素级监督优化几何,显著提升细粒度人体细节重建精度。为提升实用性,我们进一步提出模块化数据标注策略,丰富训练数据并增强模型鲁棒性。总体而言,PEAR是无需预处理的框架,可同时以超过100 FPS的速度推断EHM-s(SMPLX与缩放版FLAME)参数。在多个基准数据集上的大量实验表明,相比以往基于SMPLX的方法,本方法在姿态估计精度上取得显著提升。

原文摘要 · Abstract (English)

Reconstructing detailed 3D human meshes from a single in-the-wild image remains a fundamental challenge in computer vision. Existing SMPLX-based methods often suffer from slow inference, produce only coarse body poses, and exhibit misalignments or unnatural artifacts in fine-grained regions such as the face and hands. These issues make current approaches difficult to apply to downstream tasks. To address these challenges, we propose PEAR-a fast and robust framework for pixel-aligned expressive human mesh recovery. PEAR explicitly tackles three major limitations of existing methods: slow inference, inaccurate localization of fine-grained human pose details, and insufficient facial expression capture. Specifically, to enable real-time SMPLX parameter inference, we depart from prior designs that rely on high resolution inputs or multi-branch architectures. Instead, we adopt a clean and unified ViT-based model capable of recovering coarse 3D human geometry. To compensate for the loss of fine-grained details caused by this simplified architecture, we introduce pixel-level supervision to optimize the geometry, significantly improving the reconstruction accuracy of fine-grained human details. To make this approach practical, we further propose a modular data annotation strategy that enriches the training data and enhances the robustness of the model. Overall, PEAR is a preprocessing-free framework that can simultaneously infer EHM-s (SMPLX and scaled-FLAME) parameters at over 100 FPS. Extensive experiments on multiple benchmark datasets demonstrate that our method achieves substantial improvements in pose estimation accuracy compared to previous SMPLX-based approaches. Project page: https://wujh2001.github.io/PEAR

3D人体重建扩散模型实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。