arXiv:2609.06221cs.RO2026-09

解决机器人语言操作中的指代混淆问题,提升复杂场景下动作准确性。

RefGuard: Identity-Aware Language-Guided Robot Manipulation via Joint Target-Anchor-Frame Grounding

论文配图:RefGuard: Identity-Aware Language-Guided Robot Manipulation via Joint Target-Anchor-Frame Grounding
图 1 · 摘自论文原文
  • 延迟对目标、锚点、参考帧的单独判断,联合保持置信度
  • 真实机器人测试中零指代错误,正确执行率达90.0%
  • 适合高精度需求场景,尤其在物体重复或空间描述模糊时

视觉-语言-动作(VLA)模型显著推动了语言引导的机器人操作,但可靠执行仍依赖于准确识别指令所指的物理对象。在包含重复物体、模糊锚点或依赖参考系的空间术语的杂乱场景中,机器人可能对语义匹配但非目标的对象执行几何上正确的动作;我们称此为身份切换。指代由目标、锚点和参考帧三个耦合隐变量共同决定,提前确定任一变量会将残余歧义转化为无声且不可逆的错误。我们提出RefGuard,一种身份感知的联合定位框架,通过维持对三者联合后验来延迟承诺。RefGuard从RGB-D观测构建帧条件的物体中心场景图,分离帧无关几何与方向关系,并通过决策策略执行、澄清、重观测或终止。在真实UF850机械臂上,RefGuard在所有歧义压力测试中无身份切换,正确执行率达90.0%,而微调后的VLA和基于大语言模型(LLM)的基线在相同测试中身份切换比例为33%-46%;在无歧义场景中成功率保持93.3%,在执行后场景变化时可恢复86.7%。在3200次程序任务套件中,相比先确定锚点和参考帧的消融版本,其正确执行率从56.6%提升至80.5%,且延迟次数更少(19.5% vs. 43.4%)。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have substantially advanced language-guided robot manipulation, yet reliable execution still hinges on identifying which physical object an instruction refers to. In cluttered scenes containing repeated objects, ambiguous anchors, or frame-dependent spatial terms, a robot can execute a geometrically valid action on a semantically compatible but unintended instance; we call this failure an identity switch. The referent is jointly determined by three coupled latent variables: the target, the anchor, and the reference frame, so committing to any one of them before execution turns residual ambiguity into a silent and irreversible error. We propose RefGuard, an identity-aware grounding framework that delays commitment by maintaining a joint posterior over all three variables. RefGuard builds a frame-conditioned object-centric scene graph from RGB-D observations, separating frame-independent geometry from directional relations, and routes the posterior through a decision policy that executes, clarifies, reobserves, or aborts. On a real UF850 arm, RefGuard records no identity switch on any ambiguity-stress trial and executes correctly on 90.0% of them, whereas fine-tuned VLA and LLM (Large Language Model)-based baselines switch identity in 33-46% of the same trials, while retaining 93.3% success on unambiguous scenes and recovering from post-grounding scene changes in 86.7% of trials. On a 3200-episode procedural suite, it raises correct execution on solvable instructions from 56.6% to 80.5% over the ablation that commits to the anchor and frame before the target, while deferring less often (19.5% vs. 43.4%).

机器人操作语言导航身份感知多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。