arXiv:2511.19236cs.ROcs.AI2025-11被引 9

端到端语言动作模型实现人形机器人全身稳定控制

SENTINEL: A Fully End-to-End Language-Action Model for Humanoid Whole Body Control

  • 直接从语言指令和本体感知生成低层动作,无中间表示
  • 在仿真和真实机器人上均实现稳定执行,支持多模态输入转换
  • 基于流匹配生成动作块,适合复杂任务与现实部署

现有类人机器人控制系统通常依赖遥操作或分模块生成流程,前者完全由人工驱动,后者在语言理解与物理行为间缺乏紧密对齐。本文提出 SENTINEL,一种面向类人机器人全身控制的全端到端语言-动作模型。我们通过预训练全身控制器在仿真中追踪人类运动,并结合文本标注构建大规模数据集。该模型直接将语言指令与本体感知输入映射为低层动作,无需中间表示。动作块采用流匹配生成,可经残差动作头进一步优化以用于真实部署。方法在仿真和真实机器人上均展现出强语义理解能力与稳定执行性能,且可通过将输入转为文本支持多模态扩展。

原文摘要 · Abstract (English)

Existing humanoid control systems often rely on teleoperation or modular generation pipelines that separate language understanding from physical execution. However, the former is entirely human-driven, and the latter lacks tight alignment between language commands and physical behaviors. In this paper, we present SENTINEL, a fully end-to-end language-action model for humanoid whole-body control. We construct a large-scale dataset by tracking human motions in simulation using a pretrained whole body controller, combined with their text annotations. The model directly maps language commands and proprioceptive inputs to low-level actions without any intermediate representation. The model generates action chunks using flow matching, which can be subsequently refined by a residual action head for real-world deployment. Our method exhibits strong semantic understanding and stable execution on humanoid robots in both simulation and real-world deployment, and also supports multi-modal extensions by converting inputs into texts.

人形机器人端到端语言动作流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。