让机器人通过视觉语言模型实现空间感知的精准动作执行
EgoActor: Grounding Task Planning into Spatial-aware Egocentric Actions for Humanoid Robots via Visual-Language Models
- 用视觉语言模型直接将指令转为带空间意识的机器人动作
- 8B和4B模型均能在1秒内完成流畅动作推理
- 适合需要复杂环境交互的人形机器人研发者
在真实场景部署人形机器人面临巨大挑战,需在信息不全和动态环境中紧密融合感知、行走与操作,并稳健切换不同类型的子任务。为此,我们提出新任务EgoActing,即直接将高层指令转化为精确、空间感知的机器人动作。我们进一步构建EgoActor,一个统一且可扩展的视觉语言模型(VLM),能实时预测行走、转向、侧移、高度变化等运动基元、头部动作、操作命令及人机交互行为,协调感知与执行。该模型利用来自真实世界示范、空间推理问答和模拟环境示范的广泛监督信号,在仅含RGB的视角数据上训练,使模型在8B和4B参数规模下均能实现鲁棒、上下文感知的决策与快速动作推断(<1秒)。在模拟与真实环境中的大量评估表明,EgoActor有效连接抽象任务规划与具体动作执行,具备跨任务和未见环境的泛化能力。
原文摘要 · Abstract (English)
Deploying humanoid robots in real-world settings is fundamentally challenging, as it demands tight integration of perception, locomotion, and manipulation under partial-information observations and dynamically changing environments. As well as transitioning robustly between sub-tasks of different types. Towards addressing these challenges, we propose a novel task - EgoActing, which requires directly grounding high-level instructions into various, precise, spatially aware humanoid actions. We further instantiate this task by introducing EgoActor, a unified and scalable vision-language model (VLM) that can predict locomotion primitives (e.g., walk, turn, move sideways, change height), head movements, manipulation commands, and human-robot interactions to coordinate perception and execution in real-time. We leverage broad supervision over egocentric RGB-only data from real-world demonstrations, spatial reasoning question-answering, and simulated environment demonstrations, enabling EgoActor to make robust, context-aware decisions and perform fluent action inference (under 1s) with both 8B and 4B parameter models. Extensive evaluations in both simulated and real-world environments demonstrate that EgoActor effectively bridges abstract task planning and concrete motor execution, while generalizing across diverse tasks and unseen environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。