arXiv:2603.13355cs.CV2026-03中稿 · as an IEEE TVCG pa…

用头手动作和场景几何预测3D意图区域,提升混合现实交互体验。

Int3DNet: Scene-Motion Cross Attention Network for 3D Intention Prediction in Mixed Reality

  • 通过交叉注意力融合稀疏运动信号与场景点云,直接推断空间意图。
  • 在1500毫秒内保持稳定性能,跨场景泛化能力优于基线方法。
  • 适用于无需物体识别的实时交互系统,如智能问答与主动响应。

我们提出Int3DNet,一种基于场景感知的网络,直接从场景几何与头手运动线索中预测3D意图区域,无需显式物体级感知即可实现鲁棒的人类意图预测。在混合现实(MR)中,意图预测至关重要,可使系统提前预判用户行为,减少交互延迟,保障无缝用户体验。该方法采用稀疏运动线索与场景点云的交叉注意力融合机制,创新性地在场景中直接解析用户的三维空间意图。我们在MoGaze和CIRCLE数据集上进行了评估,这两个是公开的全身人-场景交互数据集,结果表明模型在长达1500毫秒的时间范围内表现一致,且在多样及未见场景中均超越基线方法。此外,我们通过基于意图区域的高效视觉问答(VQA)演示了该方法的可用性。Int3DNet能够从头手运动和场景几何中可靠生成3D意图区域,从而通过主动处理意图区域,实现人与MR系统间的流畅交互。

原文摘要 · Abstract (English)

We propose Int3DNet, a scene-aware network that predicts 3D intention areas directly from scene geometry and head-hand motion cues, enabling robust human intention prediction without explicit object-level perception. In Mixed Reality (MR), intention prediction is critical as it enables the system to anticipate user actions and respond proactively, reducing interaction delays and ensuring seamless user experiences. Our method employs a cross attention fusion of sparse motion cues and scene point clouds, offering a novel approach that directly interprets the user's spatial intention within the scene. We evaluated Int3DNet on MoGaze and CIRCLE datasets, which are public datasets for full-body human-scene interactions, showing consistent performance across time horizons of up to 1500 ms and outperforming the baselines, even in diverse and unseen scenes. Moreover, we demonstrate the usability of proposed method through a demonstration of efficient visual question answering (VQA) based on intention areas. Int3DNet provides reliable 3D intention areas derived from head-hand motion and scene geometry, thus enabling seamless interaction between humans and MR systems through proactive processing of intention areas.

意图预测混合现实3D理解注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。