arXiv:2604.20246cs.ROcs.AI2026-04

让机器人从反应式操作转向规划式决策,提升工业场景下长时任务的可靠性。

Cortex 2.0: Grounding World Models in Real-World Industrial Deployment

论文配图:Cortex 2.0: Grounding World Models in Real-World Industrial Deployment
图 1 · 摘自论文原文
  • 在视觉隐空间生成多个未来动作轨迹,评估后选最优执行
  • 四类复杂任务中均超越现有最先进模型,成功率显著提升
  • 适合高混乱、强遮挡、需频繁接触的工业真实环境

工业机器人操作需要在不同机械臂、任务和物体分布变化下实现可靠的长时程执行。尽管视觉-语言-动作模型展现出强大泛化能力,但其本质上仍为反应式控制:仅基于当前观测优化下一步动作,无法评估潜在未来,导致长时任务中失败累积。Cortex 2.0 转向“计划-执行”范式,在视觉隐空间生成候选未来轨迹,评估其预期成功度与效率,仅选择得分最高的轨迹执行。我们在单臂与双臂操作平台上,对四类递增复杂度的任务进行了评估:拾取放置、物品与垃圾分拣、螺丝分拣、鞋盒开箱。Cortex 2.0 在所有任务中持续优于当前最先进的视觉-语言-动作基线,表现最佳。系统在充满杂乱、频繁遮挡和高接触频率的非结构化环境中依然可靠,而传统反应式策略在此类场景中失效。结果表明,基于世界模型的规划可在复杂工业环境中稳定运行。

原文摘要 · Abstract (English)

Industrial robotic manipulation demands reliable long-horizon execution across embodiments, tasks, and changing object distributions. While Vision-Language-Action models have demonstrated strong generalization, they remain fundamentally reactive. By optimizing the next action given the current observation without evaluating potential futures, they are brittle to the compounding failure modes of long-horizon tasks. Cortex 2.0 shifts from reactive control to plan-and-act by generating candidate future trajectories in visual latent space, scoring them for expected success and efficiency, then committing only to the highest-scoring candidate. We evaluate Cortex 2.0 on a single-arm and dual-arm manipulation platform across four tasks of increasing complexity: pick and place, item and trash sorting, screw sorting, and shoebox unpacking. Cortex 2.0 consistently outperforms state-of-the-art Vision-Language-Action baselines, achieving the best results across all tasks. The system remains reliable in unstructured environments characterized by heavy clutter, frequent occlusions, and contact-rich manipulation, where reactive policies fail. These results demonstrate that world-model-based planning can operate reliably in complex industrial environments.

机器人操作世界模型规划工业部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。