通过跨部位交互提升上半身姿态与形状估计精度
Chatting about Upper-Body Expressive Human Pose and Shape Estimation

- 设计协同依赖变换器,实现面部、手部与躯干的特征级互促
- 在公开数据集上达到最新最佳性能,野拍图像泛化能力强
- 专为上半身表达性姿态估计设计,适合虚拟人建模应用
表达性人体姿态与形状估计(EHPS)在AR/VR等应用中至关重要,近年来进展显著。然而,现有最先进方法在面部与手部区域参数估计仍不准确,且对野外图像泛化能力有限。为此,我们提出CoEvoer——一种面向上半身EHPS的一阶段协同交叉依赖变换器框架。该框架在不同身体部位间实现显式特征级交互,通过上下文信息交换实现相互增强:躯干等大而易估区域提供全局语义与位置先验,指导面部与手部等精细区域的估计;而面部与手部捕捉的局部细节又能反向优化相邻部位。据我们所知,CoEvoer是首个专为上半身EHPS设计的框架,旨在通过联合参数回归捕捉面部、手部与躯干间的强耦合与语义依赖。大量实验表明,CoEvoer在上半身基准测试中取得最优性能,并在未见野拍图像上展现出强大泛化能力。
原文摘要 · Abstract (English)
Expressive Human Pose and Shape Estimation (EHPS) plays a crucial role in various AR/VR applications and has witnessed significant progress in recent years. However, current state-of-the-art methods still struggle with accurate parameter estimation for facial and hand regions and exhibit limited generalization to wild images. To address these challenges, we present CoEvoer, a novel one-stage synergistic cross-dependency transformer framework tailored for upper-body EHPS. CoEvoer enables explicit feature-level interaction across different body parts, allowing for mutual enhancement through contextual information exchange. Specifically, larger and more easily estimated regions such as the torso provide global semantics and positional priors to guide the estimation of finer, more complex regions like the face and hands. Conversely, the localized details captured in facial and hand regions help refine and calibrate adjacent body parts. To the best of our knowledge, CoEvoer is the first framework designed specifically for upper-body EHPS, with the goal of capturing the strong coupling and semantic dependencies among the face, hands, and torso through joint parameter regression. Extensive experiments demonstrate that CoEvoer achieves state-of-the-art performance on upper-body benchmarks and exhibits strong generalization capability even on unseen wild images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。