arXiv:2605.30326cs.ROcs.AI2026-05

新基准RoboWits测试机器人在意外情况下的创意解题能力。

RoboWits: Unexpected Challenges for Robotic Creative Problem Solving

论文配图:RoboWits: Unexpected Challenges for Robotic Creative Problem Solving
图 1 · 摘自论文原文
  • 用多智能体框架自动生成带突变的复杂任务
  • 预训练视觉语言模型在变异任务上表现明显下降
  • 适合研究具身智能、机器人推理与鲁棒性的人看

机器人在真实环境中应对意外挑战,需具备推理、适应与创造性解决问题的能力。然而,现有机器人评测主要关注技能执行,难以评估认知推理能力。我们提出RoboWits,一个双臂机器人基准,用于系统评估认知推理、创意工具使用及对意外条件的鲁棒性。为实现可扩展的高质量推理型突发场景构建,我们设计了基于多智能体协作的自动化任务生成管道,包含种子任务生成与验证、度量生成、场景生成和任务突变等模块。利用该管道,我们整理出30个多样化种子任务,并衍生出208个具有变异与难度分级的任务,涵盖几何、材料与装配类推理。我们测试了主流机器人策略、预训练视觉语言动作模型(VLAs)和理想状态规划器。结果揭示显著性能差距:尽管预训练VLAs在单任务微调后对种子任务有一定成功,但在变异任务上表现急剧下降,表明其在需要推理、策略适应及应对欺骗或受限环境的操纵任务中存在脆弱性。项目页面见https://umass-embodied-agi.github.io/RoboWits。

原文摘要 · Abstract (English)

The ability to reason, adapt, and creatively solve problems under unexpected challenges is essential for robots operating in real-world environments. However, current robotic benchmarks primarily emphasize skill-level execution and provide limited insight into such cognitive reasoning capabilities. We introduce RoboWits, a bi-manual robotic benchmark designed to systematically evaluate cognitive reasoning, creative tool use, and robustness to unexpected conditions. To enable scalable construction of high-quality reasoning-centric unexpected scenarios, we propose an automated task generation pipeline formulated as a multi-agent cooperative framework, comprising agents for seed task generation and verification, metric generation, scene generation, and task mutation. Using the pipeline, we curated 30 diverse seed tasks and 208 tasks with mutations and graded difficulty across geometry, material, and assembly-based reasoning. We benchmark popular robot policies, pre-trained VLAs, and oracle-state planners. Our results reveal a significant performance gap: while pre-trained VLAs exhibit preliminary success on seed tasks after single-task fine-tuning, they struggle to perform on mutated tasks, implying their brittleness in manipulation tasks requiring reasoning, strategy adaptation, and robustness to deceptive or constrained environments. Project page is available at https://umass-embodied-agi.github.io/RoboWits.

机器人推理具身智能任务生成鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。