arXiv:2502.18373cs.CVcs.AI2025-02NeurIPS被引 16

构建多视角体感相机仿真与真实数据集,助力下肢动作识别

EgoSim: An Egocentric Multi-view Simulator and Real Dataset for Body-worn Cameras during Motion and Activity

  • 基于真实动作捕捉数据生成多角度体感影像,还原运动伪影
  • 包含119小时虚拟数据+5小时真实数据,覆盖6个佩戴位置
  • 适合研究可穿戴设备视觉、3D姿态估计及人体动作分析的学者

计算机视觉中的体感任务研究主要集中在头戴式摄像头,如鱼眼相机或沉浸式头显内置相机。随着光学传感器微型化,摄像头将广泛集成于身体各部位的可穿戴设备中,为人体运动追踪、姿态估计和动作识别等任务带来新视角,尤其利于解决下肢常被遮挡的问题。本文提出EgoSim,一个新型体感相机仿真器,可从穿戴者身体多个位置生成逼真的体感渲染图像。其核心特征是利用真实动作捕捉数据生成运动伪影,这对臂部或腿部佩戴摄像头尤为显著。此外,我们引入MultiEgoView数据集,包含六台体感相机拍摄的119小时虚拟数据(源自AMASS动作序列,4个高保真虚拟环境)以及5小时真实世界数据(13名参与者使用六台GoPro相机采集,配合Xsens动作捕捉系统获取全身体三维姿态)。通过训练端到端视频仅输入的3D姿态估计网络,验证了EgoSim的有效性。域差距分析表明,该数据集与仿真器显著提升真实数据推理性能。

原文摘要 · Abstract (English)

Research on egocentric tasks in computer vision has mostly focused on head-mounted cameras, such as fisheye cameras or embedded cameras inside immersive headsets. We argue that the increasing miniaturization of optical sensors will lead to the prolific integration of cameras into many more body-worn devices at various locations. This will bring fresh perspectives to established tasks in computer vision and benefit key areas such as human motion tracking, body pose estimation, or action recognition -- particularly for the lower body, which is typically occluded. In this paper, we introduce EgoSim, a novel simulator of body-worn cameras that generates realistic egocentric renderings from multiple perspectives across a wearer's body. A key feature of EgoSim is its use of real motion capture data to render motion artifacts, which are especially noticeable with arm- or leg-worn cameras. In addition, we introduce MultiEgoView, a dataset of egocentric footage from six body-worn cameras and ground-truth full-body 3D poses during several activities: 119 hours of data are derived from AMASS motion sequences in four high-fidelity virtual environments, which we augment with 5 hours of real-world motion data from 13 participants using six GoPro cameras and 3D body pose references from an Xsens motion capture suit. We demonstrate EgoSim's effectiveness by training an end-to-end video-only 3D pose estimation network. Analyzing its domain gap, we show that our dataset and simulator substantially aid training for inference on real-world data. EgoSim code & MultiEgoView dataset: https://siplab.org/projects/EgoSim

体感视觉动作捕捉3D姿态估计多视角仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。