arXiv:2602.22056cs.ROcs.LG2026-02中稿 · IROS 2026被引 4

让机器人在执行中通过少量人工修正快速自适应,避免生成错误。

FlowCorrect: Efficient Interactive Correction of Generative Flow Policies for Robotic Manipulation

  • 用轻量VR界面接收人类微调指令,实时修正动作策略。
  • 仅需少量修正即实现80%失败案例的成功率提升。
  • 无需重训练,适合真实场景下人机协同的机器人操作。

生成式操作策略在部署时遭遇分布偏移可能灾难性失败,但许多失败实为近似成功:机械臂已接近正确姿态,只需微小调整即可完成。我们提出FlowCorrect,一种模块化交互式模仿学习方法,可在不重新训练的前提下,从稀疏、相对的人工修正中实时适应流匹配操作策略。执行期间,人类通过轻量级VR界面提供短暂的姿态修正指令。FlowCorrect利用这些稀疏修正局部调整策略,在不损害原有任务性能的同时优化动作表现。我们在真实机器人上评估了四个桌面任务:抓取放置、倾倒、杯子扶正和插入。在极低修正成本下,FlowCorrect使先前失败案例的成功率提升至80%,同时保持对已解决场景的性能。结果表明,该方法能从极少示范中学习,实现快速、高效、增量式的部署阶段人机协同修正,适用于现实机器人视觉运动策略的在线优化。

原文摘要 · Abstract (English)

Generative manipulation policies can fail catastrophically under deployment-time distribution shift, yet many failures are near-misses: the robot reaches almost-correct poses and would succeed with a small corrective motion. We propose FlowCorrect, a modular interactive imitation learning approach that enables deployment-time adaptation of flow-matching manipulation policies from sparse, relative human corrections without retraining. During execution, a human provides brief corrective pose nudges via a lightweight VR interface. FlowCorrect uses these sparse corrections to locally adapt the policy, improving actions without retraining the backbone while preserving the model performance on previously learned scenarios. We evaluate on a real-world robot across four tabletop tasks: pick-and-place, pouring, cup uprighting, and insertion. With a low correction budget, FlowCorrect achieves an 80% success rate on previously failed cases while preserving performance on previously solved scenarios. The results clearly demonstrate that FlowCorrect learns from very few demonstrations and enables fast, sample-efficient, incremental, human-in-the-loop corrections of generative visuomotor policies at deployment time in real-world robotics.

机器人操作交互学习流匹配人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。