arXiv:2511.10276cs.ROcs.AI2025-11被引 2

测试通用机器人在超市环境中的表现,发现现有模型仍不足够通用。

RoboBenchMart: Benchmarking Robots in Retail Environment

  • 构建模拟零售场景,评估机器人在复杂货架上的操作能力。
  • 模型在常见购物任务中表现不佳,说明跨域泛化仍有差距。
  • 开源完整工具链,支持后续研究与模型优化。

现有机器人操作基准多集中于桌面或家庭场景,尽管推动了显著进展,但尚不清楚这些通用视觉语言动作模型(VLAs)是否能真正泛化到几何结构、语义和工作流程不同的新领域。本文提出RoboBenchMart,一个面向零售暗仓环境的开源仿真基准,要求移动操作机器人对多种生鲜商品执行复杂操作任务。该场景面临密集物品堆积、空间布局多样等挑战,商品位置高度、深度各异且紧密相邻。通过生成的轨迹数据,我们构建了当前通用视觉语言动作模型的标准微调设置,并评估多个顶尖模型。结果表明,即使在常见零售任务上,这些模型仍表现不佳,说明其跨域泛化能力尚未成熟。为促进后续研究,我们发布了RoboBenchMart套件,包含程序化店铺布局生成器、轨迹生成管道、评估工具及微调基线模型。

原文摘要 · Abstract (English)

Most existing robotic manipulation benchmarks focus on tabletop or household scenarios. While these setups have driven impressive progress, it remains unclear whether generalist VLAs that excel there can truly generalize to domains with different geometry, semantics, and workflows. We introduce RoboBenchMart, an open-source simulated benchmark targeting retail dark-store environments, where a mobile manipulator must perform complex manipulation tasks with diverse grocery items. This setting presents significant challenges, including dense object clutter and varied spatial configurations, with items positioned at different heights, depths, and in close proximity. By targeting on the retail domain, our benchmark addresses a setting with strong potential for near-term automation impact. Using generated trajectories, we model a standard, realistic fine-tuning setup for current generalist VLAs and evaluate several state-of-the-art models. We find that they still struggle even on common retail tasks, indicating that these models are not yet truly general across domains. To support further research, we release the RoboBenchMart suite, which includes a procedural store layout generator, a trajectory generation pipeline, evaluation tools, and fine-tuned baseline models.

机器人仿真基准零售自动化视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。