针对第一人称视角鱼眼图像,提出新型3D人体网格恢复模型
Fish2Mesh Transformer: 3D Human Mesh Recovery from Egocentric Vision
- 设计鱼眼感知的Transformer架构,用专属位置编码减少畸变
- 在自建数据集上实现优于现有方法的3D人体网格重建精度
- 适合研究可穿戴设备、第一人称视觉与人体姿态估计的学者
第一人称视角人体估计旨在从可穿戴摄像头的第一人称视角推断用户的身体姿态与形状。尽管已有研究利用姿态估计技术缓解头戴鱼眼镜头带来的自遮挡和图像畸变问题,但针对3D人体网格恢复(HMR)的相应进展仍有限。本文提出Fish2Mesh,一种基于Transformer的鱼眼感知模型,专用于第一人称视角下的3D人体网格恢复。我们设计了第一人称位置嵌入模块,为Swin Transformer生成专属位置表,以减轻鱼眼图像畸变影响。模型采用多任务头部进行SMPL参数回归和相机位移估计,并以3D/2D关节点作为辅助损失项辅助训练。为解决第一人称数据稀缺问题,我们利用预训练的4D-Human模型与第三人称摄像头构建弱监督训练数据集。实验表明,Fish2Mesh在多个指标上超越现有最先进3D HMR模型。
原文摘要 · Abstract (English)
Egocentric human body estimation allows for the inference of user body pose and shape from a wearable camera's first-person perspective. Although research has used pose estimation techniques to overcome self-occlusions and image distortions caused by head-mounted fisheye images, similar advances in 3D human mesh recovery (HMR) techniques have been limited. We introduce Fish2Mesh, a fisheye-aware transformer-based model designed for 3D egocentric human mesh recovery. We propose an egocentric position embedding block to generate an ego-specific position table for the Swin Transformer to reduce fisheye image distortion. Our model utilizes multi-task heads for SMPL parametric regression and camera translations, estimating 3D and 2D joints as auxiliary loss to support model training. To address the scarcity of egocentric camera data, we create a training dataset by employing the pre-trained 4D-Human model and third-person cameras for weak supervision. Our experiments demonstrate that Fish2Mesh outperforms previous state-of-the-art 3D HMR models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。