arXiv:2512.08580cs.ROcs.AI2025-12被引 7

让机器人像人一样思考并行动,实现意图驱动的智能控制。

Mind to Hand: Purposeful Robotic Control via Embodied Reasoning

  • 用三阶段预训练融合视觉、语言与动作,提升机器人推理与执行能力。
  • 在Astribot S1上实测,长周期任务和自然指令理解性能超越基线。
  • 适合研究通用机器人控制、具身智能与人机协作的开发者参考。

人类行为依赖情境与意图,推理起核心作用。尽管互联网规模数据使AI具备广泛推理能力,但将其落地到物理动作仍是重大挑战。本文提出Lumo-1,一种统一机器人推理(“心智”)与动作(“手”)的通用视觉-语言-动作(VLA)模型。方法基于预训练视觉-语言模型(VLMs)的多模态推理能力,逐步扩展至具身推理与动作预测,并最终实现结构化推理与推理-动作对齐。整个过程包含三个阶段:(1)在精选视觉-语言数据上继续预训练VLM,增强规划、空间理解与轨迹预测等具身推理技能;(2)联合使用跨具身机器人数据与视觉-语言数据进行协同训练;(3)在具备类人灵巧性与敏捷性的双臂移动操作机器人Astribot S1上收集轨迹,进行动作训练。最后引入强化学习,进一步优化推理-动作一致性,实现语义推断与运动控制的闭环。大量实验表明,Lumo-1在具身视觉-语言推理任务中表现显著提升,真实世界评估显示其在多种高难度任务中超越强基线,对新物体与环境具有强泛化能力,尤其擅长处理长周期任务及需策略、概念与空间推理的人类自然指令。

原文摘要 · Abstract (English)

Humans act with context and intention, with reasoning playing a central role. While internet-scale data has enabled broad reasoning capabilities in AI systems, grounding these abilities in physical action remains a major challenge. We introduce Lumo-1, a generalist vision-language-action (VLA) model that unifies robot reasoning ("mind") with robot action ("hand"). Our approach builds upon the general multi-modal reasoning capabilities of pre-trained vision-language models (VLMs), progressively extending them to embodied reasoning and action prediction, and ultimately towards structured reasoning and reasoning-action alignment. This results in a three-stage pre-training pipeline: (1) Continued VLM pre-training on curated vision-language data to enhance embodied reasoning skills such as planning, spatial understanding, and trajectory prediction; (2) Co-training on cross-embodiment robot data alongside vision-language data; and (3) Action training with reasoning process on trajectories collected on Astribot S1, a bimanual mobile manipulator with human-like dexterity and agility. Finally, we integrate reinforcement learning to further refine reasoning-action consistency and close the loop between semantic inference and motor control. Extensive experiments demonstrate that Lumo-1 achieves significant performance improvements in embodied vision-language reasoning, a critical component for generalist robotic control. Real-world evaluations further show that Lumo-1 surpasses strong baselines across a wide range of challenging robotic tasks, with strong generalization to novel objects and environments, excelling particularly in long-horizon tasks and responding to human-natural instructions that require reasoning over strategy, concepts and space.

具身智能机器人控制视觉语言动作长程推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。