让机器人像人一样看、听、动,实现自主交互。
Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
- 用语言与动作对齐学习通用运动模式,打通语义理解与控制
- 通过视觉上下文感知生成适应环境的动作,提升任务适应性
- 自监督生成问答数据,高效利用海量未标注视频
当前类人机器人控制框架多依赖反应式机制,因数据稀缺缺乏自主交互能力。本文提出Humanoid-VLA框架,融合语言理解、第一视角场景感知与运动控制,实现通用类人控制。该框架首先基于非第一视角人类动作数据与文本描述进行语言-动作预对齐,学习通用运动模式与动作语义;随后通过参数高效的视频条件微调引入第一视角视觉上下文,实现上下文感知的动作生成;进一步提出自监督数据增强策略,直接从运动数据生成伪标注的问答对,有效利用大规模无标签视频数据。基于全身体控架构,实验表明Humanoid-VLA在物体交互与环境探索任务中展现出更强的上下文感知能力,表现出更接近人类的自适应智能行为。
原文摘要 · Abstract (English)
This paper addresses the limitations of current humanoid robot control frameworks, which primarily rely on reactive mechanisms and lack autonomous interaction capabilities due to data scarcity. We propose Humanoid-VLA, a novel framework that integrates language understanding, egocentric scene perception, and motion control, enabling universal humanoid control. Humanoid-VLA begins with language-motion pre-alignment using non-egocentric human motion datasets paired with textual descriptions, allowing the model to learn universal motion patterns and action semantics. We then incorporate egocentric visual context through a parameter efficient video-conditioned fine-tuning, enabling context-aware motion generation. Furthermore, we introduce a self-supervised data augmentation strategy that automatically generates pseudoannotations directly derived from motion data. This process converts raw motion sequences into informative question-answer pairs, facilitating the effective use of large-scale unlabeled video data. Built upon whole-body control architectures, extensive experiments show that Humanoid-VLA achieves object interaction and environment exploration tasks with enhanced contextual awareness, demonstrating a more human-like capacity for adaptive and intelligent engagement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。