首个同步双侧肌电与视角视觉的手部姿态数据集,助力精准手部动作捕捉。
EgoEMG: A Multimodal Egocentric Dataset with Bilateral EMG and Vision for Hand Pose Estimation

- 采集41人双手动作的双腕肌电(16通道)与全景视频数据
- 在60类手势上实现超过10小时记录,支持跨手势/用户泛化测试
- 提供三类任务基准,验证肌电+视觉融合优于单一模态
表面肌电(sEMG)可记录手部运动时的肌肉活动,并解码出精细的手指关节运动。肌电与第一人称视觉在手部感知中具有互补性:肌电可在遮挡或光照差条件下捕捉细粒度指节动作,而视觉提供整体手部姿态。然而,现有数据集未同步这两类模态。本文提出EgoEMG,一个用于双手姿态估计的多模态第一人称数据集。包含双侧腕带肌电(共16通道,每侧8个,采样率2 kHz)、120 Hz IMU、第一人称广角RGB视频、外部RGB-D视频,以及基于动捕的肢体运动数据(含腕关节角度)。数据覆盖41名参与者完成60类手势(30类单手、30类双手),总时长超10小时。我们还构建了三项基准任务——肌电到姿态、视觉到姿态、肌电+视觉融合,采用统一的关节角预测目标和共享的泛化分割策略(跨手势、跨用户、联合)。作为基线,我们评估了EMGFormer(肌电到姿态)及通用ResNet/ViT骨干网络(视觉到姿态)。进一步研究了一种残差融合架构,性能优于等效轻量级纯视觉模型。EgoEMG及其基准为未来多模态手部姿态估计研究奠定基础。
原文摘要 · Abstract (English)
Surface electromyography (sEMG) records muscle activity during hand movement and can be decoded to recover detailed hand articulation. EMG and egocentric vision are complementary for hand sensing: EMG captures fine-grained finger articulation even under occlusion and poor lighting, while vision provides global hand configuration. However, no existing dataset synchronizes both modalities. We present EgoEMG, a multimodal egocentric dataset for bimanual hand pose estimation. EgoEMG includes bilateral wristband EMG with 16 total channels (8 per wrist) sampled at 2 kHz, 120 Hz IMU, egocentric wide-angle RGB video, external RGB-D video, and mocap-derived hand motion with wrist articulation angles. The dataset covers 41 participants performing 60 gesture classes, including 30 single-hand gestures and 30 bimanual gestures, totaling more than 10 hours of recording. We also introduce a benchmark with three tasks -- EMG-to-pose, vision-to-pose, and EMG+vision fusion -- under a shared joint-angle prediction target and common generalization split axes (cross-gesture, cross-user, and combined). As baselines, we evaluate EMGFormer for EMG-to-pose and generic ResNet/ViT backbones for vision-to-pose. We further study a residual fusion architecture that improves over matched lightweight vision-only baselines. Together, EgoEMG and its benchmark establish a foundation for future research on multimodal hand pose estimation with EMG and vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。