arXiv:2606.09286cs.RO2026-06

让人形机器人在复杂环境中自主完成搬运、推车等高难度动作

VAIC: Vision-Guided Humanoid Agile Object Interaction Control via Decoupled Commands

论文配图:VAIC: Vision-Guided Humanoid Agile Object Interaction Control via Decoupled Commands
图 1 · 摘自论文原文
  • 用分步指令和视觉反馈替代精确轨迹,提升部署灵活性
  • 单个策略成功执行搬运、推车、滑板等多种动态任务
  • 适合希望实现真实场景下人形机器人自主操作的研究者

人形机器人在现实世界中辅助潜力巨大,但要在非结构化环境中实现敏捷物体交互,需全身高度协同。当前控制器存在关键部署瓶颈:依赖密集参考轨迹与完美状态观测,限制了物理泛化能力。本文提出视觉引导的敏捷交互控制框架VAIC,仅使用机载深度信息、历史本体感知及解耦用户指令接口。采用两阶段蒸馏机制:先由具备特权信息的教师策略学习多样化交互技能;再由可部署的学生策略通过多轴速度目标与交互指示符替代全身体追踪,并结合循环物体适应模块,从原始深度流与本体感知中隐式推断不可观测的物体动态。在人形机器人上的评估与真实部署表明,单一VAIC策略能稳定完成箱体搬运、推车互动、滑板等多样动态任务,显著优于基线,推动自主人形机器人实用化进展。

原文摘要 · Abstract (English)

Humanoid robots hold immense potential for real-world assistance, yet agile interaction with objects in unstructured environments demands tightly coupled whole-body coordination. Despite recent advancements, current controllers face a critical deployment gap. They rely heavily on dense reference trajectories and perfect state observability, which inherently limits physical generalization. We present Vision Guided Agile Interaction Control (VAIC), a unified framework that bridges this gap by operating exclusively on onboard depth, historical proprioception, and a decoupled user command interface. VAIC employs a two-stage distillation paradigm. First, a privileged teacher policy masters diverse interaction skills using precise object kinematics and exact environmental states. Second, a deployable student policy distills these capabilities by replacing full body tracking with velocity targets across multiple axes and an interaction indicator for each frame. The student utilizes a recurrent object adaptation module to implicitly infer unobservable object dynamics from raw depth streams and proprioception. Evaluations and real-world deployments on the humanoid robot demonstrate that a single VAIC policy successfully executes highly diverse dynamic tasks. These tasks include box carrying, cart interaction, and skateboarding, consistently outperforming baselines and advancing autonomous humanoid deployment.

人形机器人视觉控制动态交互强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。