arXiv:2603.21051cs.RO2026-03被引 1

受大脑启发的双流视觉模型,提升机器人操作的空间与动态适应能力。

Cortical Policy: A Dual-Stream View Transformer for Robotic Manipulation

  • 双流架构分别处理静态与动态视觉信息,模拟人类脑部处理机制。
  • 在RLBench和COLOSSEUM上超越现有方法,实现在复杂场景下的高成功率。
  • 适合需要强空间推理与实时适应的机器人视觉控制任务。

视图变换器通过多视角观测预测动作,在机器人操作中表现优异。现有方法通常以视图特异性方式提取静态视觉表征,导致三维空间推理能力不足且缺乏动态适应性。受人类大脑整合静态与动态视图的启发,我们提出Cortical Policy,一种新型双流视图变换器,同时从静态视图与动态视图流中联合推理。静态视图流通过预训练3D基础模型提取的几何一致关键点对齐特征,增强空间理解;动态视图流通过自中心注视估计模型的位置感知预训练,计算上复现人类皮层背侧通路,实现自适应调整。随后,两流互补的视图表征被融合以决定最终动作,使模型能在语言指令下处理空间复杂且动态变化的任务。在RLBench、挑战性的COLOSSEUM基准及真实世界任务上的实验表明,Cortical Policy显著优于现有最先进基线,验证了双流设计在视觉-运动控制中的优越性。这一受大脑启发的框架为机器人操作提供了新视角,并有望拓展至更广泛的基于视觉的机器人控制应用。

原文摘要 · Abstract (English)

View transformers process multi-view observations to predict actions and have shown impressive performance in robotic manipulation. Existing methods typically extract static visual representations in a view-specific manner, leading to inadequate 3D spatial reasoning ability and a lack of dynamic adaptation. Taking inspiration from how the human brain integrates static and dynamic views to address these challenges, we propose Cortical Policy, a novel dual-stream view transformer for robotic manipulation that jointly reasons from static-view and dynamic-view streams. The static-view stream enhances spatial understanding by aligning features of geometrically consistent keypoints extracted from a pretrained 3D foundation model. The dynamic-view stream achieves adaptive adjustment through position-aware pretraining of an egocentric gaze estimation model, computationally replicating the human cortical dorsal pathway. Subsequently, the complementary view representations of both streams are integrated to determine the final actions, enabling the model to handle spatially-complex and dynamically-changing tasks under language conditions. Empirical evaluations on RLBench, the challenging COLOSSEUM benchmark, and real-world tasks demonstrate that Cortical Policy outperforms state-of-the-art baselines substantially, validating the superiority of dual-stream design for visuomotor control. Our cortex-inspired framework offers a fresh perspective for robotic manipulation and holds potential for broader application in vision-based robot control.

机器人操作双流模型视觉控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。