arXiv:2607.09776cs.ROcs.AI2026-07

通过角色分工提升人机协作效率,实现机器人大规模后训练。

HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation

论文配图:HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation
图 1 · 摘自论文原文
  • 分工两类操作员:远程干预与现场监控,减少任务切换
  • 支持12台机器人并发,成功率80%~95%,吞吐量提升1.7~4.2倍
  • 自动生成轨迹分段,精准识别失败模式,适合工业级机器人部署

在将视觉语言动作(VLA)模型适配至下游任务时,常需多次迭代后训练以修复策略缺陷。本文提出HELP,一种人类高效的机器人大规模后训练流程,由两名专业化操作员协同管理十二台机器人。一名训练过的遥操作员提供高价值远程干预与恢复示范,另一名场地操作员负责监控机器人集群、触发接管并执行物理复位。角色分工减少任务切换,降低培训成本,扩大交互覆盖范围。并发监督不仅提升数据采集量,还丰富策略行为观测,便于发现重复失败模式,实现更精准的干预与复位。为高效利用大规模混合质量回放数据,HELP引入VLAC,一种专为该场景设计的自动轨迹分段批判器,可将自主轨迹分为进展型、空转、故障诱导和恢复四类。有效片段与人机协同数据结合用于下一轮后训练。在四项真实世界操作任务中,HELP达成80%–95%成功率,任务吞吐量相较基线提升1.7×–4.2×。在相同人机协同回收预算下,VLAC-CUT进一步使吞吐量增益达1.20×–3.43×,成功率增益达1.50×–3.00×。

原文摘要 · Abstract (English)

When adapting Vision Language Action (VLA) models to downstream tasks, multiple rounds of post-training are often required to progressively address policy weaknesses. In this report, we focus on maximizing human efficiency during this iterative process, measured by policy improvement and task throughput per unit of human labor and time. We propose HELP, a Human-Efficient Large-scale robot Post-training pipeline in which two specialized operators supervise twelve robots concurrently. A trained Teleoperator provides high-value remote interventions and recovery demonstrations, while a Floor Operator monitors the robot fleet, triggers takeovers, and performs physical resets. This role specialization improves human efficiency by reducing task switching, lowering operator training costs, and expanding robot interaction coverage. Beyond increasing rollout volume, concurrent supervision also broadens the range of policy behaviors observed by the human team, making recurring failure modes easier to identify and enabling more targeted takeovers, resets, and recovery demonstrations. To efficiently utilize the large and mixed-quality rollout data, HELP incorporates \vlac, an automatic rollout segmentation critic specifically designed for this setting. It separates autonomous trajectories into progress-making, idle, failure-inducing, and recovery segments. Useful rollout segments are retained and combined with Human-in-the-Loop data for the next post-training round. Across four real-world manipulation tasks, HELP achieves 80\%--95\% success rates and improves task throughput by 1.7$\times$--4.2$\times$ over the base model. Under matched HITL recovery budgets, VLAC-CUT further amplifies throughput gains by 1.20$\times$--3.43$\times$ and success-rate gains by 1.50$\times$--3.00$\times$ over HITL-only updates.

机器人人机协同后训练轨迹分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。