用触觉信号同时预测动作、姿态和进度,提升人机交互性能
Shared Representation for 3D Pose Estimation, Action Classification, and Progress Prediction from Tactile Signals
- 设计共享模型统一处理三类任务,利用多任务学习增强表征
- 在15人、8类活动、7小时数据上实现三项任务性能超越现有方法
- 首个基于足部触觉信号的动作进度预测研究,适合人机交互场景
估计人体姿态、分类动作并预测运动进展是人机交互中的关键任务。视觉方法在真实环境中易受遮挡和隐私问题影响,而触觉传感可避免这些问题。然而,以往的触觉方法均独立处理各项任务,导致性能不佳。本文提出一种共享卷积-变换器模型SCOTTI,通过学习统一表征,同步完成三维人体姿态估计、动作分类与动作完成进度预测。据我们所知,这是首个利用定制无线鞋垫传感器的足部触觉信号进行动作进度预测的研究。该联合框架借助多任务学习的协同优势,在所有任务上均优于独立训练的模型。实验表明,SCOTTI在三个任务中均表现更优。此外,本文还构建了一个新数据集,涵盖15名参与者执行8种不同活动,总时长7小时。
原文摘要 · Abstract (English)
Estimating human pose, classifying actions, and predicting movement progress are essential for human-robot interaction. While vision-based methods suffer from occlusion and privacy concerns in realistic environments, tactile sensing avoids these issues. However, prior tactile-based approaches handle each task separately, leading to suboptimal performance. In this study, we propose a Shared COnvolutional Transformer for Tactile Inference (SCOTTI) that learns a shared representation to simultaneously address three separate prediction tasks: 3D human pose estimation, action class categorization, and action completion progress estimation. To the best of our knowledge, this is the first work to explore action progress prediction using foot tactile signals from custom wireless insole sensors. This unified approach leverages the mutual benefits of multi-task learning, enabling the model to achieve improved performance across all three tasks compared to learning them independently. Experimental results demonstrate that SCOTTI outperforms existing approaches across all three tasks. Additionally, we introduce a novel dataset collected from 15 participants performing various activities and exercises, with 7 hours of total duration, across eight different activities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。