用视觉强化学习让模拟人形机器人自发产生搜寻和手眼协调行为。
Emergent Active Perception and Dexterity of Simulated Humanoids from Visual Reinforcement Learning
- 仅用第一视角视觉输入,通过强化学习训练单一策略完成多任务
- 从零训练可自然涌现主动搜索等类人行为,无需预设状态信息
- 适合研究动画、机器人控制与具身智能中的感知-动作闭环
人类行为从根本上由视觉感知塑造——我们与世界互动的能力依赖于主动获取相关信息并相应调整动作。像寻找物体、伸手和手眼协调这样的行为,自然源于感官系统的结构。受此启发,我们提出感知灵巧控制(Perceptive Dexterous Control, PDC),一种基于视觉驱动的模拟人形机器人全身灵巧控制框架。PDC仅依赖第一视角视觉进行任务设定,通过视觉线索实现物体搜索、目标放置和技能选择,不依赖特权状态信息(如3D物体位置与几何)。这种感知即接口范式使单一策略能够执行多种家庭任务,包括伸手、抓取、放置及复杂物体操作。我们还表明,从零开始的强化学习训练可催生主动搜索等涌现行为。结果表明,视觉驱动控制与复杂任务能诱发类人行为,是实现动画、机器人与具身智能中感知-动作闭环的关键要素。
原文摘要 · Abstract (English)
Human behavior is fundamentally shaped by visual perception -- our ability to interact with the world depends on actively gathering relevant information and adapting our movements accordingly. Behaviors like searching for objects, reaching, and hand-eye coordination naturally emerge from the structure of our sensory system. Inspired by these principles, we introduce Perceptive Dexterous Control (PDC), a framework for vision-driven dexterous whole-body control with simulated humanoids. PDC operates solely on egocentric vision for task specification, enabling object search, target placement, and skill selection through visual cues, without relying on privileged state information (e.g., 3D object positions and geometries). This perception-as-interface paradigm enables learning a single policy to perform multiple household tasks, including reaching, grasping, placing, and articulated object manipulation. We also show that training from scratch with reinforcement learning can produce emergent behaviors such as active search. These results demonstrate how vision-driven control and complex tasks induce human-like behaviors and can serve as the key ingredients in closing the perception-action loop for animation, robotics, and embodied AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。