arXiv:2605.20072cs.AIcs.RO2026-05被引 2

高精度视觉输入反而降低机器人语言模型解谜能力,噪声有帮助。

Probing Embodied LLMs: When Higher Observation Fidelity Hurts Problem Solving

论文配图:Probing Embodied LLMs: When Higher Observation Fidelity Hurts Problem Solving
图 1 · 摘自论文原文
  • 用不同感知信息测试机器人语言模型解谜行为
  • 40%感知错误时成功率提升2.85倍,最佳表现
  • 适合关注智能体感知与推理交互的研究者

大型语言模型正被越来越多地用于机器人系统作为认知组件,但其决策过程不透明,难以解释闭环具身任务中的成功或失败。基于实证人工智能方法,我们通过改变智能体可用信息并测量行为变化,来行为化地研究具身语言模型的表现。在物理机器人平台上,使用包含隐藏依赖关系的机械锁盒任务,评估了模型在RGB、RGB-D和真实符号观测下的表现,并通过受控仿真探究其行为。出人意料的是,模型在原始RGB输入下表现最佳,在理想真值观测下最差。在仿真中,随机翻转动作结果后发现,适度噪声(40%翻转概率)能提升性能,成功率相较无噪声基线提高2.85倍。进一步分析表明,该增益源于减少重复动作循环。这些发现表明,仅看成功率不足以评估语言模型,因为表现可能反映感知误差与推理失败之间的交互,而非真正的鲁棒解题能力。

原文摘要 · Abstract (English)

Large Language Models are increasingly proposed as cognitive components for robotic systems, yet their opaque decision processes make it difficult to explain success or failure in closed-loop embodied tasks. Following an empirical AI methodology, we study embodied LLM agents behaviorally by varying the information available to the agent and measuring the resulting changes in behavior. Using the Lockbox, a sequential mechanical puzzle with hidden interdependencies, we evaluate LLMs across RGB, RGB-D, and ground-truth symbolic observations in a physical robotic setup and use controlled simulation to probe the resulting behavior. Counterintuitively, agents perform best under raw RGB input and worst under perfect ground-truth observations. In simulation, we probe this effect by randomly flipping perceived action outcomes and find that moderate noise improves performance, peaking at a 40% flip probability with a 2.85-fold success rate increase over the noise-free baseline. Further analysis links this gain to a reduction in repetitive action loops. These findings suggest that success rates alone are insufficient for evaluating LLMs, as measured performance may reflect the interaction between perceptual errors and reasoning failures rather than robust problem solving.

具身智能语言模型机器人感知噪声

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。