为机器人视角下的自然人机交互构建新数据集与长距离动作识别方法
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
- 提出多模态机器人视角动作数据集ACTIVE,覆盖30类动作、80人、46868段视频
- 在3-50米远距离下实现高精度动作识别,突破传统方法在移动平台上的局限
- 适合研究服务机器人感知、长距离动作识别及多模态融合的科研人员
自然人机交互(N-HRI)要求机器人在不同距离和状态下识别人类动作,无论自身是否运动。现有基准因数据量少、模态单一、类别有限及场景多样性不足,难以应对N-HRI的独特挑战。为此,我们提出ACTIVE(Action from Robotic View)——一个专为移动服务机器人常见视觉视角设计的大规模数据集。ACTIVE包含30种复合动作类别、80名参与者、46,868个标注视频实例,涵盖RGB与点云双模态。参与者在多种环境中执行动作,距离范围为3米至50米,摄像头平台同步移动,模拟真实地面不平导致的相机高度变化。该数据集旨在推动N-HRI中动作与属性识别的研究。此外,我们提出ACTIVE-PC方法,采用多层邻域采样、分层识别器、弹性椭圆查询及运动干扰解耦技术,实现远距离精准动作感知。实验验证了其有效性。代码已开源。
原文摘要 · Abstract (English)
Natural Human-Robot Interaction (N-HRI) requires robots to recognize human actions at varying distances and states, regardless of whether the robot itself is in motion or stationary. This setup is more flexible and practical than conventional human action recognition tasks. However, existing benchmarks designed for traditional action recognition fail to address the unique complexities in N-HRI due to limited data, modalities, task categories, and diversity of subjects and environments. To address these challenges, we introduce ACTIVE (Action from Robotic View), a large-scale dataset tailored specifically for perception-centric robotic views prevalent in mobile service robots. ACTIVE comprises 30 composite action categories, 80 participants, and 46,868 annotated video instances, covering both RGB and point cloud modalities. Participants performed various human actions in diverse environments at distances ranging from 3m to 50m, while the camera platform was also mobile, simulating real-world scenarios of robot perception with varying camera heights due to uneven ground. This comprehensive and challenging benchmark aims to advance action and attribute recognition research in N-HRI. Furthermore, we propose ACTIVE-PC, a method that accurately perceives human actions at long distances using Multilevel Neighborhood Sampling, Layered Recognizers, Elastic Ellipse Query, and precise decoupling of kinematic interference from human actions. Experimental results demonstrate the effectiveness of ACTIVE-PC. Our code is available at: https://github.com/wangzy01/ACTIVE-Action-from-Robotic-View.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。