arXiv:2607.05377cs.ROcs.AI2026-07被引 1

让机器人长时任务执行更智能,通过双向对齐规划与动作

Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation

论文配图:Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation
图 1 · 摘自论文原文
  • 用32种标准技能动作连接高层规划与底层执行
  • 自动生成4000小时视频+30小时仿真数据,提升任务成功率
  • 零样本完成复杂真实任务,适合通用机器人研发

尽管近期视觉-语言-动作(VLA)模型在通用操作中展现出潜力,但其马尔可夫特性导致难以处理长时序任务。现有分层双系统方法虽缓解此问题,却存在高层语义与底层运动之间的鸿沟。本文提出Cortex框架,通过定制化规划接口,实现高层视觉语言模型(VLM)到低层视觉语言动作模型(VLA)的双向对齐,将操作子任务标准化为32种通用技能原语,并在数据生成中注入可执行性原则,如代表性物体属性和轨迹可达性优化。该方法实现了超过4000小时开源视频数据的自动标注及30小时仿真数据生成。进一步设计事件均衡采样策略,构建用于微调训练的数据集,以更好应对子任务切换中的规划歧义;推理阶段通过任务上下文到技能约束的精心工程设计增强鲁棒性。开环VLM与闭环系统评估均验证其有效性:在Libero-long上优于单体基线3.1%,在RoboTwin上提升4.1%。尤为突出的是,基于通用VLM的Cortex可零样本完成未见的真实长时任务,如多阶段化学实验,仅需搭配微调后的VLA即可实现——这是仅靠VLA微调无法达成的能力。

原文摘要 · Abstract (English)

While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature-relying solely on current observations. Hierarchical dual-system methods address this but suffer from a gap between high-level planning semantics and low-level execution kinematics. We introduce Cortex, a bidirectionally aligned embodied agent framework with a customized planning interface that conveys executable and tractable subtask plans from high-level VLM to low-level VLA. Specifically, we standardize manipulation subtasks into 32 canonical skill primitives and inject tractability principles, such as representative object attributes and improved trajectory reachability, into the data generation pipeline. This enables automatic annotation of over 4k hours of open-source video data and generation of 30 hours of simulation data. We further devise an event-balanced sampling strategy to construct training data for fine-tuning the framework to better handle planning ambiguity during subtask transitions, enhanced by carefully designed harness engineering from task contexts to skill constraints during inference. Both open-loop VLM and closed-loop system evaluations demonstrate Cortex's efficacy, e.g., it outperforms monolithic baselines by 3.1% on Libero-long and 4.1% on RoboTwin. Notably, Cortex's generalist VLM enables zero-shot completion of unseen real-world long-horizon tasks, such as multi-stage chemistry experiments, by simply combining with a fine-tuned VLA-a capability infeasible through VLA fine-tuning alone.

机器人控制长时任务多模态通用智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。