用眼神和手势预测手部动作,助力康复机器人理解患者意图。
Gaze-Guided 3D Hand Motion Prediction for Detecting Intent in Egocentric Grasping Tasks
- 融合视线、历史手势和环境物体信息动态预测手部运动。
- 仅需少量输入帧即可准确预测,眼神信息显著提升效果。
- 适合神经康复场景,无需预先知道要抓的物体。
在神经康复应用中,通过手部动作预测来检测人类意图对驱动上肢辅助机器人至关重要。传统依赖生理信号的方法受限且常缺乏环境上下文。本文提出一种新方法,可同时预测未来的手部姿态与关节位置序列。该方法融合了视线信息、历史手部运动序列及环境物体数据,能在不事先知道目标抓取物的情况下,自适应满足患者的辅助需求。具体采用向量量化变分自编码器进行鲁棒的手部姿态编码,并结合自回归生成式变换器实现高效的手部运动序列预测。我们在健康受试者中开展试点研究验证了该技术的可行性。为训练与评估,我们收集了来自多名受试者的多种抓取动作数据集。大量实验表明,所提方法能成功预测连续手部动作;尤其在输入帧数较少时,视线信息显著提升预测能力,展现了其在真实场景中的潜力。
原文摘要 · Abstract (English)
Human intention detection with hand motion prediction is critical to drive the upper-extremity assistive robots in neurorehabilitation applications. However, the traditional methods relying on physiological signal measurement are restrictive and often lack environmental context. We propose a novel approach that predicts future sequences of both hand poses and joint positions. This method integrates gaze information, historical hand motion sequences, and environmental object data, adapting dynamically to the assistive needs of the patient without prior knowledge of the intended object for grasping. Specifically, we use a vector-quantized variational autoencoder for robust hand pose encoding with an autoregressive generative transformer for effective hand motion sequence prediction. We demonstrate the usability of these novel techniques in a pilot study with healthy subjects. To train and evaluate the proposed method, we collect a dataset consisting of various types of grasp actions on different objects from multiple subjects. Through extensive experiments, we demonstrate that the proposed method can successfully predict sequential hand movement. Especially, the gaze information shows significant enhancements in prediction capabilities, particularly with fewer input frames, highlighting the potential of the proposed method for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。