机器人通过双阶段推断人类意图,实现更稳定协作。
Probabilistic Human Intent Prediction for Mobile Manipulation: An Evaluation with Human-Inspired Constraints
- 分导航与操作两阶段,用双重信念层追踪目标
- 导航稳定性达93%-100%,操作稳定率超94%
- 提前23.6秒识别操作目标,适合人机协同场景
准确推断人类意图可实现无需限制人类控制的人机协作。本文提出GUIDER(全局用户意图双阶段估计框架),使机器人能实时估计操作者意图。该框架包含两个耦合的信念层:分别追踪导航与操作目标。导航阶段通过协同图融合控制器速度与占用网格,对交互区域排序;到达目标后,自主多视角扫描构建局部3D点云。操作阶段结合U2Net显著性、FastSAM实例显著性及三项几何抓取可行性测试,并采用末端执行器运动学感知更新规则,动态演化物体概率。GUIDER可在无预设目标情况下识别意图区域与物体。在Isaac Sim中进行25次实验(五名参与者×五种任务变体),对比两种基线模型。导航阶段,其稳定性为93%-100%(基线60%-100%),在重定向任务(T5)中提升39.5%;操作阶段稳定性达94%-100%(基线69%-100%),重定向任务(T3)中差值为31.4%。在几何约束试验中,其识别操作意图时间比Trajectron早3倍(中位剩余预测时间23.6秒 vs 7.8秒)。结果验证了双阶段框架有效性,显著提升移动操作中的意图推断能力。
原文摘要 · Abstract (English)
Accurate inference of human intent enables human-robot collaboration without constraining human control or causing conflicts between humans and robots. We present GUIDER (Global User Intent Dual-phase Estimation for Robots), a probabilistic framework that enables a robot to estimate the intent of human operators. GUIDER maintains two coupled belief layers, one tracking navigation goals and the other manipulation goals. In the Navigation phase, a Synergy Map blends controller velocity with an occupancy grid to rank interaction areas. Upon arrival at a goal, an autonomous multi-view scan builds a local 3D cloud. The Manipulation phase combines U2Net saliency, FastSAM instance saliency, and three geometric grasp-feasibility tests, with an end-effector kinematics-aware update rule that evolves object probabilities in real-time. GUIDER can recognize areas and objects of intent without predefined goals. We evaluated GUIDER on 25 trials (five participants x five task variants) in Isaac Sim, and compared it with two baselines, one for navigation and one for manipulation. Across the 25 trials, GUIDER achieved a median stability of 93-100% during navigation, compared with 60-100% for the BOIR baseline, with an improvement of 39.5% in a redirection scenario (T5). During manipulation, stability reached 94-100% (versus 69-100% for Trajectron), with a 31.4% difference in a redirection task (T3). In geometry-constrained trials (manipulation), GUIDER recognized the object intent three times earlier than Trajectron (median remaining time to confident prediction 23.6 s vs 7.8 s). These results validate our dual-phase framework and show improvements in intent inference in both phases of mobile manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。