用VR运动数据识别用户决策中的困惑、犹豫和准备状态,准确率达82%。
Cognitive State Inference from VR Motion via Motion Foundation Model
- 通过适配器将稀疏VR数据映射到预训练动作模型空间,实现跨用户迁移。
- 在24人小数据集上,基于大模型的分类准确率达82%,媲美人类观察者。
- 首次将动作基础模型用于VR认知状态推断,适合元宇宙与交互研究者。
随着虚拟现实(VR)普及,消费级设备采集的头手运动数据日益常见。本文探究仅凭此类遥测数据能否识别决策过程中的瞬时认知状态——包括困惑、犹豫和准备状态。我们构建了一个新数据集,包含结构化决策任务中采集的头手运动,并带有逐帧标注。在两种评估协议下测试经典机器学习模型、时序神经网络及动作基础模型:(1) 同一用户的未来预测;(2) 跨用户泛化至未见用户。提出一种针对VR的运动适配器,将稀疏的VR遥测映射为与大规模全身动作预训练模型兼容的表示,无需显式重建完整动作。据我们所知,这是首个将动作基础模型适配至VR运动以完成分类任务的工作。结果表明,仅靠运动信号即可捕捉有意义的认知特征,且预训练模型即使在小样本(24名参与者)下也优于传统方法。本方法达到82%准确率,与人类观察者相当甚至超越。研究揭示了VR运动蕴含比以往预期更丰富的行为信息,凸显大规模动作预训练在扩展现实(XR)应用中的潜力。我们将公开数据集与建模框架以支持后续研究。
原文摘要 · Abstract (English)
As virtual reality (VR) becomes widespread, head and hand motion data captured by consumer systems has become substantially more common. However, the extent of what can be inferred from such motion remains unclear. This paper investigates whether transient cognitive states, specifically confusion, hesitation, and readiness during different stages of decision-making, can be inferred from VR telemetry alone. We introduce a novel dataset of head and hand motion collected during structured decision-making tasks, with frame-level annotations of these states. We evaluate classical machine learning models, temporal neural networks, and motion foundation models under two protocols: (1) future-in-time prediction for the same users, and (2) cross-user generalization to unseen users. We further propose a VR-native motion adapter that maps sparse VR telemetry to representations compatible with motion foundation models pretrained on large-scale full-body motion data, enabling transfer without explicit full-body reconstruction. To our knowledge, this is the first work to adapt a motion foundation model to VR motion for a classification task. Results show that motion-only sensing captures meaningful signals of cognitive states, and that pretrained motion foundation models generalize more effectively than classical and temporal models even with a small dataset of 24 participants. Our approach achieves 82% accuracy, comparable to and sometimes surpassing human observers. These findings suggest that VR motion encodes richer behavioral information than previously assumed and highlight the potential of large-scale motion pretraining for XR applications. We will release the dataset and modeling framework to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。