arXiv:2510.18267cs.CVcs.AI2025-10中稿 · ICME2025

通过低维交互优化,提升复杂场景下人体网格重建精度与效率

Latent-Info and Low-Dimensional Learning for Human Mesh Recovery and Parallel Optimization

  • 分两阶段利用频域特征挖掘人体形状与动作的潜在信息
  • 在保持高精度的同时,计算量显著低于现有注意力方法
  • 适合需要高效高精度3D人体重建的应用场景

现有3D人体网格恢复方法常未能充分挖掘潜在信息(如人体动作、形状对齐),导致肢体错位及局部细节不足,尤其在复杂场景中。同时,使用注意力机制建模网格顶点与姿态节点交互虽提升性能,但计算成本高昂。为此,本文提出一种基于潜在信息与低维学习的两阶段网络。第一阶段从图像特征的高低频成分中全面提取全局(如整体形状对齐)与局部(如纹理、细节)信息,并聚合为混合潜在频域特征,有效挖掘潜在信息;第二阶段借助该特征,通过降维与并行优化实现粗略3D人体模板与3D姿态间的交互学习,优化姿态与形状。相比现有方法,本方案在降低计算开销的同时维持重建精度。大量公开数据集上的实验表明,其性能优于当前最先进方法。

原文摘要 · Abstract (English)

Existing 3D human mesh recovery methods often fail to fully exploit the latent information (e.g., human motion, shape alignment), leading to issues with limb misalignment and insufficient local details in the reconstructed human mesh (especially in complex scenes). Furthermore, the performance improvement gained by modelling mesh vertices and pose node interactions using attention mechanisms comes at a high computational cost. To address these issues, we propose a two-stage network for human mesh recovery based on latent information and low dimensional learning. Specifically, the first stage of the network fully excavates global (e.g., the overall shape alignment) and local (e.g., textures, detail) information from the low and high-frequency components of image features and aggregates this information into a hybrid latent frequency domain feature. This strategy effectively extracts latent information. Subsequently, utilizing extracted hybrid latent frequency domain features collaborates to enhance 2D poses to 3D learning. In the second stage, with the assistance of hybrid latent features, we model the interaction learning between the rough 3D human mesh template and the 3D pose, optimizing the pose and shape of the human mesh. Unlike existing mesh pose interaction methods, we design a low-dimensional mesh pose interaction method through dimensionality reduction and parallel optimization that significantly reduces computational costs without sacrificing reconstruction accuracy. Extensive experimental results on large publicly available datasets indicate superiority compared to the most state-of-the-art.

3D人体重建低维学习网格优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。