让人形机器人在复杂环境中智能抓取动态物体,无需外部感知。
HAIC: Humanoid Agile Object Interaction Control via Dynamics-Aware World Model
- 用自身体感历史预测物体运动状态,构建动态占位图。
- 在多种负载下成功完成滑板、推拉车等敏捷任务,成功率高。
- 适合研究人形机器人交互控制与多物体长期任务的学者。
人形机器人在非结构化环境中执行复杂全身任务潜力巨大。尽管人-物交互(HOI)已取得进展,但多数方法聚焦于与机器人刚性耦合的全驱动物体,忽略了具有独立动力学和非完整约束的欠驱动物体,后者引入了耦合力和遮挡带来的控制挑战。本文提出HAIC,一种无需外部状态估计即可实现跨多种物体动力学鲁棒交互的统一框架。核心贡献是动力学预测器,仅通过本体感受历史估计物体的高阶状态(速度、加速度)。这些预测投影到静态几何先验上,形成空间对齐的动态占位图,使策略能推断盲区中的碰撞边界与接触可能性。采用非对称微调机制,世界模型持续适应学生策略的探索行为,确保分布偏移下的鲁棒状态估计。实验表明,该框架在人形机器人上实现了高成功率,可主动补偿惯性扰动,在不同负载下完成滑板、推拉车等敏捷任务;同时掌握多物体长时序任务,如在复杂地形上搬运箱子,通过预测多个物体的动力学实现协同控制。
原文摘要 · Abstract (English)
Humanoid robots show promise for complex whole-body tasks in unstructured environments. Although Human-Object Interaction (HOI) has advanced, most methods focus on fully actuated objects rigidly coupled to the robot, ignoring underactuated objects with independent dynamics and non-holonomic constraints. These introduce control challenges from coupling forces and occlusions. We present HAIC, a unified framework for robust interaction across diverse object dynamics without external state estimation. Our key contribution is a dynamics predictor that estimates high-order object states (velocity, acceleration) solely from proprioceptive history. These predictions are projected onto static geometric priors to form a spatially grounded dynamic occupancy map, enabling the policy to infer collision boundaries and contact affordances in blind spots. We use asymmetric fine-tuning, where a world model continuously adapts to the student policy's exploration, ensuring robust state estimation under distribution shifts. Experiments on a humanoid robot show HAIC achieves high success rates in agile tasks (skateboarding, cart pushing/pulling under various loads) by proactively compensating for inertial perturbations, and also masters multi-object long-horizon tasks like carrying a box across varied terrain by predicting the dynamics of multiple objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。