arXiv:2607.15641cs.ROcs.AI2026-07中稿 · SemRob Workshop, R…

IMBench评测机器人在复杂物理任务中的直觉式操作能力。

IMBench: A Benchmark for Intuitive Robotic Manipulation

论文配图:IMBench: A Benchmark for Intuitive Robotic Manipulation
图 1 · 摘自论文原文
  • 设计35个任务,融合感知、推理、动作生成与迭代执行
  • 现有模型在满足约束和泛化上表现不足,差距明显
  • 适合研究具身智能、机器人通用策略的学者使用

人类通过将推理与运动控制结合,在多种约束下解决复杂的操作任务,具备对物理世界的理解能力,能将思维转化为行动并快速适应新场景、新任务和新规则。我们称这种能力为直觉操作。现有基准未能捕捉这一整合性:要么孤立评估物理推理,要么仅衡量策略性能而不要求显式推理。我们提出IMBench,一个旨在评估从感知到执行全过程整合能力的基准,涵盖感知、物理推理、动作生成与迭代执行。任务要求模型推断任务相关的物理结构,并在明确约束下生成可行的动作序列,包括接触密集型操作、工具使用及多阶段依赖。该基准包含35个任务和14,000条过滤后的轨迹,并提供可扩展的工具以生成多样化场景。实验显示显著差距:视觉语言模型展现出部分物理推理能力,但无法生成可执行计划;最先进的视觉-语言-动作模型则难以满足任务约束且泛化能力弱。这些结果揭示当前基础模型与通用机器人策略在直觉操作方面存在缺失,表明IMBench是迈向评估与提升更整合、自适应物理智能的重要一步。

原文摘要 · Abstract (English)

Humans combine reasoning and motor control to solve complex manipulation tasks under diverse constraints. They build an understanding of the physical world that helps them convert reasoning into actions and quickly adapt to new scenes, tasks, and rules. We refer to this capability as intuitive manipulation. Existing benchmarks fail to capture this integration: they evaluate physical reasoning in isolation from execution, or measure policy performance without requiring explicit reasoning. We introduce IMBENCH, a benchmark designed to evaluate intuitive manipulation as an integrated capability spanning perception, physical reasoning, action generation, and iterative execution. Our tasks require models to infer task-relevant physical structure and generate feasible action sequences under explicit constraints, including contact-rich manipulation, tool use, and multi-stage dependencies. We introduce a benchmark of 35 tasks, 14K filtered trajectories, and scalable tools for generating diverse scenarios. Experiments reveal a consistent gap: vision language models show partial physical reasoning ability but fail to produce executable plans, while state-of-the-art vision-language-action models struggle to satisfy task constraints and generalize across scenarios. These results identify intuitive manipulation as a missing axis in current foundation models and generalist robot policies, and position IMBENCH as a step toward evaluating and enabling more integrated, adaptive physical intelligence.

机器人操作物理推理基准评测具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。