让机器人通过接触环境来精准执行复杂动作
SceneBot: Contact-Prompted General Humanoid Whole Body Tracking with Scene-Interaction

- 用接触标签引导单一策略,统一处理行走与交互
- 基于7.5小时重构数据训练,实现跨场景泛化
- 适合需要精细环境互动的机器人控制研究
当前的人形机器人强化学习策略在自由空间运动中表现优异,但在涉及接触的复杂任务中表现不佳,因纯运动学追踪无法解决与物体和不平地形交互时的物理模糊性。为此,我们提出SceneBot,一个统一的运动追踪框架,可同时处理自由空间移动、地形穿越和全身操作。SceneBot将单一策略同时依赖参考动作和各关节接触标签,明确设定预期的环境交互。为解决标注交互数据匮乏问题,我们提出一种事后场景重建方法,从重定向的人体运动中推断场景交互图。在7.5小时重构的接触丰富数据上训练后,SceneBot成功泛化至未见过的动作与环境。结果表明,SceneBot是首个无缝融合自由空间与接触密集行为的通用框架,可执行如搬箱上楼等复杂长周期任务,并确立接触条件作为人形控制的强大接口。所有代码与数据将开源。更多演示与信息见:https://ericcsr.github.io/scenebot/
原文摘要 · Abstract (English)
Current humanoid reinforcement-learning policies excel at free-space motions but struggle with contact-rich tasks, as pure kinematic tracking cannot resolve the physical ambiguities of interacting with objects and uneven terrain. To address this, we introduce SceneBot, a unified motion-tracking framework capable of handling freespace locomotion, terrain traversal, and whole-body manipulation. SceneBot conditions a single policy on both reference motions and per-link contact labels, explicitly defining expected environmental interactions. To overcome the lack of annotated interaction data, we propose a hindsight scene reconstruction approach that infers scene-interaction graphs from retargeted human motion. Trained on 7.5 hours of this reconstructed, contact-rich data, SceneBot successfully generalizes to unseen motions and environments. Our results demonstrate that SceneBot is the first general framework to seamlessly unify free-space and contact-rich behaviors executing complex, long-horizon tasks like carrying a box upstairs and establishing contact conditioning as a powerful interface for humanoid control. All code and data will be open-sourced. More demos and information are available at: https://ericcsr.github.io/scenebot/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。