arXiv:2502.17399cs.HCcs.RO2025-02

用单帧视角追踪相同物体,让AR游戏在动态场景中实现稳定虚实交互。

Enriching physical-virtual interaction in AR gaming by tracking identical objects via an egocentric partial observation frame

  • 基于整数规划与沃罗诺伊图剪枝,仅凭一帧视角重识别相同物体。
  • 实验显示计算时间减少50%,准确率仍保持91%。
  • 适合无标记、动态环境下的可穿戴AR设备使用。

增强现实(AR)游戏,尤其是为头戴式显示器设计的,日益普及。然而,现有系统大多依赖预扫描的静态环境,并严重依赖连续跟踪或基于标记的方案,限制了在动态物理空间中的适应性。这对头戴式设备尤其成问题,因其随用户头部移动,无法保持对场景的固定视角。此外,持续场景观察在功耗和处理能力有限的可穿戴设备上既不节能也不实用。当环境中存在多个相同物体时,标准跟踪流程往往因缺乏持续观测或外部传感器而无法维持身份一致性。这阻碍了在动态或遮挡场景中流畅的虚实交互。为此,我们提出一种基于优化的新框架,仅通过头显捕捉的一帧局部视角,实现对相同物体的重新识别。将问题建模为标签分配任务,采用整数规划求解,并引入沃罗诺伊图剪枝策略提升计算效率。模拟实验表明,计算时间减少50%,准确率仍达91%。我们在定量合成与真实世界实验中评估了该方法,并进行了三项定性真实实验,验证其在动态、无标记物体交互中的实用性与泛化能力。视频演示见:https://youtu.be/RwptEfLtW1U。

原文摘要 · Abstract (English)

Augmented reality (AR) games, particularly those designed for head-mounted displays, have grown increasingly prevalent. However, most existing systems depend on pre-scanned, static environments and rely heavily on continuous tracking or marker-based solutions, which limit adaptability in dynamic physical spaces. This is particularly problematic for AR headsets and glasses, which typically follow the user's head movement and cannot maintain a fixed, stationary view of the scene. Moreover, continuous scene observation is neither power-efficient nor practical for wearable devices, given their limited battery and processing capabilities. A persistent challenge arises when multiple identical objects are present in the environment-standard object tracking pipelines often fail to maintain consistent identities without uninterrupted observation or external sensors. These limitations hinder fluid physical-virtual interactions, especially in dynamic or occluded scenes where continuous tracking is infeasible. To address this, we introduce a novel optimization-based framework for re-identifying identical objects in AR scenes using only one partial egocentric observation frame captured by a headset. We formulate the problem as a label assignment task solved via integer programming, augmented with a Voronoi diagram-based pruning strategy to improve computational efficiency. This method reduces computation time by 50% while preserving 91% accuracy in simulated experiments. Moreover, we evaluated our approach in quantitative synthetic and quantitative real-world experiments. We also conducted three qualitative real-world experiments to demonstrate the practical utility and generalizability for enabling dynamic, markerless object interaction in AR environments. Our video demo is available at https://youtu.be/RwptEfLtW1U.

AR游戏物体重识别可穿戴设备无标记交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。