arXiv:2606.15568cs.RO2026-06

用人类操作与预训练模型动作融合,提升机器人任务成功率。

SAPS: Shared Autonomy for Policy Steering by Blending Teleoperation with a Pretrained VLA

论文配图:SAPS: Shared Autonomy for Policy Steering by Blending Teleoperation with a Pretrained VLA
图 1 · 摘自论文原文
  • 在动作层面融合人类指令与预训练模型输出,无需重训练或改架构。
  • 仿真与真实机器人测试中任务成功率最高提升82%。
  • 减少人工干预,比纯遥控更快完成任务,适合辅助操作和数据采集。

视觉-语言-动作(VLA)模型在机器人操作中展现出强大泛化能力,但在分布外的空间或语义扰动下仍易失效。虽然人类远程操控可提供可靠恢复,但需高认知负荷和精准控制,现有策略常需额外模型或采样器修改。本文提出行动级共享自主框架SAPS,将实时人类操控指令与预训练策略动作在动作层面融合,无需策略重训练、辅助动力学模型或结构改动。我们设计并评估三种仲裁策略,包括基于余弦相似度的动态协调机制,以衡量人类与策略动作的几何一致性。在仿真(LIBERO、LIBERO-PRO、CALVIN)及真实机器人硬件上验证,SAPS相较自主执行在仿真和真实场景中任务成功率最高提升82%。同时,该方法显著降低人类干预频率,且任务完成速度优于自主执行与纯遥控。结果表明,行动级共享自主是一种模型无关、实用性强的方法,适用于现实世界中带人类操作员的通用机器人部署,具有辅助遥控与规模化数据收集的前景。

原文摘要 · Abstract (English)

Recent advancements in Vision-Language-Action (VLA) models have demonstrated impressive generalist capabilities in robot manipulation, yet these policies can be brittle under out-of-distribution spatial and semantic perturbations. While human teleoperation offers reliable recovery, it can demand high cognitive load and precise manual control, and existing policy steering methods often require auxiliary models or sampler modifications. In this work, we introduce Shared Autonomy for Policy Steering (SAPS), a framework that blends real-time human teleoperation commands with pretrained policy actions at the action level. SAPS requires no policy retraining, auxiliary dynamics models, or architectural modifications. We propose and evaluate three arbitration strategies to balance human and VLA policy control, including a dynamic Cosine-similarity arbitration strategy that computes the geometric agreement between human and policy actions. Across evaluations in simulation (LIBERO, LIBERO-PRO, CALVIN) and on real-world robot hardware, SAPS improves task success rates over autonomous execution by up to 82% in both simulation and the real world. Furthermore, our approach drastically reduces human intervention compared to pure teleoperation, while simultaneously achieving faster task completion times than both autonomous execution and pure teleoperation. These results demonstrate that action-level shared autonomy is a practical, model-agnostic approach for reliably deploying generalist robot policies in real-world contexts involving a human operator,with promising applications in assistive teleoperation and scalable data collection.

机器人共享自主视觉语言动作人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。