arXiv:2510.10742cs.CVcs.LG2025-10

预测用户在虚拟现实中的交互行为,让系统提前响应。

Seeing My Future: Predicting Situated Interaction Behavior in Virtual Reality

  • 分层建模人类意图,结合环境上下文预测具体动作。
  • 在真实世界数据集上各项指标均优于现有方法。
  • 适合开发主动适应用户的智能虚拟现实应用。

虚拟和增强现实系统日益需要根据用户行为进行智能适应以提升交互体验。这要求准确理解人类意图并预测未来的具体情境行为(如注视方向、物体交互),这对构建响应式VR/AR环境及个性化助手至关重要。然而,精准预测需建模驱动人-环境交互的潜在认知过程。本文提出一种分层、意图感知的框架,通过利用认知机制建模人类意图,并基于历史行为动态与场景上下文,识别潜在交互目标并预测细粒度未来行为。我们设计了一种动态图卷积网络(GCN)以有效捕捉人-环境关系。在挑战性真实世界基准数据集及实时VR环境中进行的大量实验表明,该方法在所有指标上均表现更优,可支持主动式VR系统实现用户行为预判与环境自适应。

原文摘要 · Abstract (English)

Virtual and augmented reality systems increasingly demand intelligent adaptation to user behaviors for enhanced interaction experiences. Achieving this requires accurately understanding human intentions and predicting future situated behaviors - such as gaze direction and object interactions - which is vital for creating responsive VR/AR environments and applications like personalized assistants. However, accurate behavioral prediction demands modeling the underlying cognitive processes that drive human-environment interactions. In this work, we introduce a hierarchical, intention-aware framework that models human intentions and predicts detailed situated behaviors by leveraging cognitive mechanisms. Given historical human dynamics and the observation of scene contexts, our framework first identifies potential interaction targets and forecasts fine-grained future behaviors. We propose a dynamic Graph Convolutional Network (GCN) to effectively capture human-environment relationships. Extensive experiments on challenging real-world benchmarks and live VR environment demonstrate the effectiveness of our approach, achieving superior performance across all metrics and enabling practical applications for proactive VR systems that anticipate user behaviors and adapt virtual environments accordingly.

虚拟现实行为预测图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。