用可变形注意力提升单图人体网格重建精度
DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery
- 采用可变形交叉注意力机制增强视觉特征捕捉能力
- 在3DPW和RICH数据集上达到单帧回归方法新纪录
- 适合关注3D人体建模与视觉解码的开发者
人体网格恢复(HMR)是计算机视觉中的关键挑战,广泛应用于动作捕捉、增强现实与生物力学。从单张图像精确预测人体姿态参数仍具难度。本文提出DeforHMR,一种基于回归的单目人体网格恢复框架,利用预训练视觉变压器(ViT)编码器提取特征,并通过新颖的查询无关可变形交叉注意力机制,在解码器中实现对空间特征的灵活、数据驱动的关注。该机制使模型能更精准地捕获局部空间信息。在3DPW和RICH两个主流3D HMR基准上,DeforHMR实现了单帧回归方法的最先进性能,推动了3D人体网格重建的发展,为从大型预训练视觉编码器中解码局部空间信息提供了新范式。
原文摘要 · Abstract (English)
Human Mesh Recovery (HMR) is an important yet challenging problem with applications across various domains including motion capture, augmented reality, and biomechanics. Accurately predicting human pose parameters from a single image remains a challenging 3D computer vision task. In this work, we introduce DeforHMR, a novel regression-based monocular HMR framework designed to enhance the prediction of human pose parameters using deformable attention transformers. DeforHMR leverages a novel query-agnostic deformable cross-attention mechanism within the transformer decoder to effectively regress the visual features extracted from a frozen pretrained vision transformer (ViT) encoder. The proposed deformable cross-attention mechanism allows the model to attend to relevant spatial features more flexibly and in a data-dependent manner. Equipped with a transformer decoder capable of spatially-nuanced attention, DeforHMR achieves state-of-the-art performance for single-frame regression-based methods on the widely used 3D HMR benchmarks 3DPW and RICH. By pushing the boundary on the field of 3D human mesh recovery through deformable attention, we introduce an new, effective paradigm for decoding local spatial information from large pretrained vision encoders in computer vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。