让机器人更懂执行:用四大能力分工的视觉语言模型
Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

- 按执行流程分四类能力,由专用模型分别训练
- 2B与35B规模模型均实现四大能力统一且损失可控
- 适合需要闭环执行的机器人任务研究与开发
视觉语言模型正成为具身智能体的推理核心。机器人执行本质上是迭代过程:每一步动作都会改变场景与物理状态,持续刷新需感知、推理和验证的内容。这要求具备不同监督信号、预测格式和验证标准的互补能力。现有方法多针对单一任务目标训练,未解决如何围绕整体执行整合这些能力。我们提出Capek 0.5,一个以执行为中心的能力分类架构的具身视觉语言模型。该模型将具身能力按执行中的功能角色分为四类:空间推理、时间理解、动作引导和状态验证。每类能力由专用专家通过强化学习获得,奖励来自共享主干网络;随后通过权重空间合并与路由策略蒸馏,整合为单个推理模型。我们在2B和35B-A3B规模上实例化Capek 0.5,从三方面评估:涵盖Capek-StateBench的新状态验证基准、能力从专家到统一体的保留率分析、以及模拟具身环境中的闭环评估。结果表明,相比初始化,多数基准性能提升,四种能力在单检查点中保持完整且损失量化,模型可成功迁移至闭环具身任务执行。
原文摘要 · Abstract (English)
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verified. Meeting these demands requires complementary capabilities that differ in supervision signals, prediction formats, and verification criteria. Existing approaches typically develop these capabilities against isolated, task-specific objectives, leaving open how they should be organized and integrated around execution as a whole. We present Capek 0.5, an embodied vision-language model built around an execution-centric capability taxonomy. Rather than organizing training by datasets or tasks, the taxonomy groups embodied capabilities according to their functional roles throughout execution and comprises four capability families: Spatial Reasoning, Temporal Understanding, Action Guidance, and State Verification. Each capability is first acquired by a dedicated specialist through reinforcement learning with verifiable rewards from a shared backbone, and the specialists are then consolidated into a single inference-time model through weight-space merging followed by routed policy-space distillation. We instantiate Capek 0.5 at the 2B and 35B-A3B scales and evaluate it from three complementary perspectives: comprehensive benchmark suites including Capek-StateBench, a new benchmark for state verification; a controlled study of capability retention from specialists to the unified model; and closed-loop evaluation in simulated embodied environments. Capek 0.5 improves the large majority of matched benchmark rows over its initialization, retains all four specialized capabilities in one checkpoint with quantified losses, and transfers to closed-loop embodied task execution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。