arXiv:2604.22491cs.HCcs.RO2026-04

通过融合指指点和抓握动作提升远距物体选择准确率

Point & Grasp: Flexible Selection of Out-of-Reach Objects Through Probabilistic Cue Integration

论文配图:Point & Grasp: Flexible Selection of Out-of-Reach Objects Through Probabilistic Cue Integration
图 1 · 摘自论文原文
  • 用概率方法融合指指点与抓握手势,灵活应对线索不可靠情况
  • 在远距抓取数据集上训练,识别出传统数据集未覆盖的抓握模式
  • 用户测试显示比单线索方法快且准,适合混合现实交互场景

在混合现实(MR)中,选择远距离物体是一项基础任务。现有方法依赖单一线索或确定性融合多线索,当主导线索失效时性能下降。本文提出一种概率线索融合框架,支持灵活组合用户生成的多种线索进行意图推断。受自然抓握行为启发,实例化为指指点与抓握手势结合的新交互方式——Point&Grasp。为此,我们构建了远距抓取(ORG)数据集,用于训练鲁棒的姿势线索似然模型,捕捉现有近距数据集中缺失的抓握模式。用户研究显示,该方法在准确性与速度上均优于单线索基线,并在多种模糊情境下仍保持实际有效性,显著优于当前最优方法。代码与数据集已开源。

原文摘要 · Abstract (English)

Selecting out-of-reach objects is a fundamental task in mixed reality (MR). Existing methods rely on a single cue or deterministically fuse multiple cues, leading to performance degradation when the dominant cue becomes unreliable. In this work, we introduce a probabilistic cue integration framework that enables flexible combination of multiple user-generated cues for intent inference. Inspired by natural grasping behavior, we instantiate the framework with pointing direction and grasp gestures as a new interaction technique, Point&Grasp. To this end, we collect the Out-of-Reach Grasping (ORG) dataset to train a robust likelihood model of the gestural cue, which captures grasping patterns not present in existing in-reach datasets. User studies demonstrate that our selection method with cue integration not only improves accuracy and speed over single-cue baselines, but also remains practically effective compared to state-of-the-art methods across various sources of ambiguity. The dataset and code are available at https://github.com/drlxj/point-and-grasp.

人机交互手势识别混合现实概率建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。