arXiv:2609.05260cs.RO2026-09

测试机器人如何根据指令微调动作,发现多条件融合是新难题。

One Word, Different Action: A Real-Robot Benchmark for Language-Conditioned Embodied Reasoning

论文配图:One Word, Different Action: A Real-Robot Benchmark for Language-Conditioned Embodied Reasoning
图 1 · 摘自论文原文
  • 用真实机器人对比相同与不同指令下的动作一致性
  • 多约束条件下模型表现显著下降,单约束已接近极限
  • 适合研究具身智能与语言理解的交叉方向

自然语言指令的变化可直接改变机器人行为。可靠的具身系统应在任务不变时保持动作一致,在任务变化时正确更新行为。我们提出「One Word, Different Action」,一个基于物理决策状态和可执行动作的真实机器人基准,通过任务保持与任务变更的指令对,联合评估决策不变性与决策敏感性,并在多约束推理和真实RGB视觉接地条件下进一步验证。实验表明,现代模型在单一约束指令变化下已接近饱和,但在需整合多个任务约束为单一可执行决策时,多个模型表现明显下降。结果表明,当前更关键的挑战不再是识别孤立指令变化,而是可靠地将多个任务要求组合成正确的机器人动作决策。

原文摘要 · Abstract (English)

Natural-language instruction changes can directly alter robot behavior. A reliable embodied system should preserve its action when the task is unchanged and update it correctly when the task itself changes. We introduce One Word, Different Action, a real-robot benchmark built on physical decision states and executable actions, using task-preserving and task-changing instruction pairs to jointly evaluate Decision Invariance and Decision Sensitivity, with further evaluation under multi-constraint reasoning and real-RGB grounding. Experiments show that modern models are near saturation on single-constraint instruction changes, yet several models degrade noticeably when multiple task constraints must be integrated into one executable decision. These results suggest that the more salient remaining challenge is no longer recognizing an isolated instruction change, but reliably composing multiple task requirements into a correct robot action decision.

具身智能语言理解机器人多约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。