用LSTM预测VR中用户抓握意图,误差小于0.25秒
Predicting User Grasp Intentions in Virtual Reality
- 基于手部运动时间序列,采用LSTM回归预测抓握意图
- 关键2秒内时间误差<0.25秒,距离误差5-20cm
- 适合做自适应触觉反馈的VR交互系统研发
在虚拟现实(VR)中预测用户意图对实现沉浸式体验至关重要,尤其在涉及复杂抓握动作且需精确触觉反馈的任务中。本文利用手部运动的时间序列数据,在810次不同物体类型、尺寸和操作方式的实验中评估了分类与回归方法的表现。结果表明,分类模型在跨用户泛化上表现不佳,性能不稳定;而基于LSTM的回归方法表现更稳健,关键两秒窗口内时间误差低于0.25秒,距离误差约为5-20厘米。尽管如此,精确预测手部姿态仍具挑战。通过分析用户差异与模型可解释性,我们探讨了部分模型失败的原因,并说明回归模型更能适应用户行为的动态复杂性。研究结果表明,机器学习模型可有效提升VR交互体验,尤其在自适应触觉反馈方面具有潜力,为实时预测用户提供重要基础。
原文摘要 · Abstract (English)
Predicting user intentions in virtual reality (VR) is crucial for creating immersive experiences, particularly in tasks involving complex grasping motions where accurate haptic feedback is essential. In this work, we leverage time-series data from hand movements to evaluate both classification and regression approaches across 810 trials with varied object types, sizes, and manipulations. Our findings reveal that classification models struggle to generalize across users, leading to inconsistent performance. In contrast, regression-based approaches, particularly those using Long Short Term Memory (LSTM) networks, demonstrate more robust performance, with timing errors within 0.25 seconds and distance errors around 5-20 cm in the critical two-second window before a grasp. Despite these improvements, predicting precise hand postures remains challenging. Through a comprehensive analysis of user variability and model interpretability, we explore why certain models fail and how regression models better accommodate the dynamic and complex nature of user behavior in VR. Our results underscore the potential of machine learning models to enhance VR interactions, particularly through adaptive haptic feedback, and lay the groundwork for future advancements in real-time prediction of user actions in VR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。