arXiv:2604.21741cs.RO2026-04被引 5

用世界模型替代真实机器人,实现高效、可复用的纠错训练。

Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training

论文配图:Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training
图 1 · 摘自论文原文
  • 用学习的世界模型代替真实机器人进行闭环推理与纠错
  • 平均提升真实任务成功率37.9点,优于基线19.0点
  • 支持状态回滚与多分支修正,适合需要密集反馈的机器人调优

后训练对将通用机器人策略转化为可靠的任务专用控制器至关重要,但现有基于人工干预的流程仍依赖真实世界执行:每次修正都需要机器人运行、场景布置、重置及操作员监督。而动作条件世界模型通常仅用于想象、合成数据生成或策略评估。本文提出「人类在世界模型中(Hi-WM)」框架,利用学习的世界模型作为可重复使用的纠错基础。先在世界模型内闭环运行策略;当推演出现错误或高风险时,人类直接在模型中提供短时修正动作。该框架缓存中间状态,支持回滚与分支,使单一失败状态可被多次用于不同修正路径,从而在基础策略表现差的区域生成密集监督信号。修正轨迹随后回加入训练集进行后训练。我们在三个涵盖刚体与柔性物体交互的真实世界操控任务上,使用两种策略骨干验证了该方法。结果显示,相比基线策略,真实世界成功率平均提升37.9点,较世界模型闭环基线提升19.0点;且世界模型评估结果与真实性能高度相关(r = 0.953)。表明世界模型不仅能生成或评估策略,还能作为可扩展的纠错平台。

原文摘要 · Abstract (English)

Post-training is essential for turning pretrained generalist robot policies into reliable task-specific controllers, but existing human-in-the-loop pipelines remain tied to physical execution: each correction requires robot time, scene setup, resets, and operator supervision in the real world. Meanwhile, action-conditioned world models have been studied mainly for imagination, synthetic data generation, and policy evaluation. We propose \textbf{Human-in-the-World-Model (Hi-WM)}, a post-training framework that uses a learned world model as a reusable corrective substrate for failure-targeted policy improvement. A policy is first rolled out in closed loop inside the world model; when the rollout becomes incorrect or failure-prone, a human intervenes directly in the model to provide short corrective actions. Hi-WM caches intermediate states and supports rollback and branching, allowing a single failure state to be reused for multiple corrective continuations and yielding dense supervision around behaviors that the base policy handles poorly. The resulting corrective trajectories are then added back to the training set for post-training. We evaluate Hi-WM on three real-world manipulation tasks spanning both rigid and deformable object interaction, and on two policy backbones. Hi-WM improves real-world success by 37.9 points on average over the base policy and by 19.0 points over a world-model closed-loop baseline, while world-model evaluation correlates strongly with real-world performance (r = 0.953). These results suggest that world models can serve not only as generators or evaluators, but also as effective corrective substrates for scalable robot post-training.

机器人后训练世界模型人机协作强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。