人类引导机器人通过偏好优化改进动作策略,实现部署后持续迭代。
Human-assisted Robotic Policy Refinement via Action Preference Optimization
- 用人类干预收集交互轨迹,构建可优化的失败修正数据
- 提出自适应重加权算法,解决不可逆动作与令牌分布不匹配问题
- 适合希望实现人机协作迭代优化的机器人研发团队
构建可靠且可迭代优化的机器人系统对实际应用至关重要。尽管视觉-语言-动作(VLA)模型被视为机器人部署的基础模型,但其依赖离线专家示范,限制了部署后的优化能力。为此,我们提出动作偏好优化(APO),通过人机交互获取偏好信号来精炼VLA模型。该方法首先建立人机协作框架,由人类干预完成可靠故障修正并收集交互轨迹。然而,直接利用这些轨迹进行偏好优化面临不可逆动作和令牌分布不匹配的挑战。为此,APO提出基于交互的二元优劣信号自适应重加权算法,有效抑制高风险动作并增强纠正动作的适应性。最终,APO使VLA模型具备从失败中学习的能力,为动态环境中持续迭代优化和可靠部署奠定基础。仿真与真实场景实验验证了该框架在多种操作任务中的优越泛化性和鲁棒性。代码与数据集已开源。
原文摘要 · Abstract (English)
Establishing a reliable and iteratively refined robotic system is essential for deploying real-world applications. While Vision-Language-Action (VLA) models are widely recognized as the foundation model for such robotic deployment, their reliance on offline expert demonstrations critically limits their capacity for post-deployment refinement. To mitigate this limitation, we introduce Action Preference Optimization (APO), a method designed to refine VLA models by human-assisted preference alignment gathered through interaction with environments. This method begins with a human-robot collaboration framework for reliable failure correction and interaction trajectory collection through human intervention. However, directly leveraging these interaction trajectories for preference optimization is non-trivial due to the challenges of irreversible robotic actions and token distribution mismatch. To solve this, APO proposes an adaptive reweighting algorithm with binary desirability signals derived from interaction, empowering VLA models effectively suppress failure-prone actions while enhancing corrective action adaptation. Ultimately, APO equips VLA models with the crucial capability to learn from failure, paving the way for their iterative refinement and reliable deployment in dynamic environments. The experiments conducted in simulation and real-world scenarios prove superior generalization and robustness of our human-assisted framework across a variety of manipulation tasks. We believe this work could bring insights for efficient and stable optimization of VLA models through human-robot collaboration. The code and dataset are released at https://github.com/GeWu-Lab/Action-Preference-Optimization
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。