arXiv:2608.17584cs.RO2026-08

让机器人像人一样实时响应任务变化,边做边改。

HODAgent: Towards On-Demand, Responsive Humanoids for Physical World Human Interaction

  • 用分层记忆与双模式架构实现边执行边调整任务
  • 仿真中成功率超基线9.8至18.9个百分点,实物测试达92%
  • 适合需要灵活应对的现实服务场景,如酒店迎宾

我们提出 HODAgent,一种面向服务场景的人形机器人系统-2级具身智能体,解决情境意图理解、响应式执行、任务修订和结果验证问题。其半双工架构集成环境交互器、规划器、执行器与分层记忆,维持服务过程中交互、规划与任务状态的一致性。该设计支持运动中接收新请求、保留进度、修改动作,并以执行结果为依据闭环确认。共享接口连接仿真与实体机器人(Unitree G1),屏蔽平台差异。在包含164个案例的交互仿真中,HODAgent 在两种视觉语言模型(VLM)后端下分别取得84.8%和91.5%的联合成功率,优于基线9.8和18.9个百分点。在物理机器人上,原子任务通过率为92%,复合任务72%,完整任务63.3%。在多个具身基准测试中,性能提升0.7至9.0分。结果表明,统一的系统-2智能体可实现仿真与现实间自适应的人形服务。

原文摘要 · Abstract (English)

We propose HODAgent, a System-2 embodied agent for humanoid robots in service settings, addressing situated intent, responsive execution, task revision, and outcome verification. Its semi-duplex architecture integrates an Env-Interactor, Planner, Executor, and hierarchical Memory to maintain coherent interaction, planning, and task state during service episodes. This allows handling new requests during motion, retaining progress, revising actions, and grounding closure in execution outcomes. A shared interface connects simulation and physical robots (Unitree G1), isolating platform-specific control. In an interactive simulation with 164 cases, HODAgent achieves 84.8% and 91.5% Joint Success under two VLM backbones, outperforming baselines by 9.8 and 18.9 points. On physical robots, pass rates are 92% (atomic), 72% (composite), and 63.3% (complete tasks). On multiple embodied benchmarks, it improves over baselines by 0.7-9.0 points. Results show a unified System-2 agent enables adaptive humanoid service across simulation and reality.

人形机器人具身智能服务场景动态任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。