用人类纠正动作指导噪声空间强化学习,提升机器人操作适应效率
UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation

- 将人类纠正动作转换为噪声目标,联合优化噪声策略
- 实测在66分钟内成功率从20%升至90%
- 适合需快速适配真实场景的机器人学习研究者
基于扩散模型的视觉-语言-动作(VLA)系统是机器人操作的强大先验,但将其适配到真实世界仍具挑战。由于在线强化学习成本高,高效适配依赖有限交互下的策略改进。噪声空间强化学习通过固定预训练VLA作为去噪生成器,仅更新轻量级噪声预测策略来降低成本,但自主探索效率低。人类干预可减轻探索负担,但其提供的是动作空间反馈,而噪声空间微调需要噪声变量监督。为此,我们提出UniSteer:一种统一噪声引导框架,通过近似动作到噪声的逆映射,将人类纠正动作转化为噪声目标,指导同一噪声策略的强化学习优化。在多个真实操作任务上的实验表明,UniSteer比现有噪声空间及动作空间人机协同基线更高效,平均在66分钟内将成功率从20%提升至90%。
原文摘要 · Abstract (English)
Diffusion-based vision-language-action (VLA) models have emerged as strong priors for robotic manipulation, yet adapting them to real-world distributions remains challenging. In particular, on-robot reinforcement learning (RL) is expensive and time-consuming, so effective adaptation depends on efficient policy improvement within a limited budget of real-world interactions. Noise-space RL lowers the cost by keeping the pretrained VLA fixed as a denoising generator while updating only a lightweight actor that predicts the noise. However, its performance is still limited due to inefficient autonomous exploration. Human corrective interventions can reduce this exploration burden, but they are naturally provided in action space, whereas noise-space finetuning requires supervision over noise variables. To address these challenges, we propose UniSteer, a Unified Noise Steering framework that combines human corrective guidance with noise-space RL through approximate action-to-noise inversion. Given a human corrective action, UniSteer inverts the frozen flow-matching decoder to recover a noise target, which provides supervised guidance for the same noise actor that is simultaneously optimized via reinforcement learning. Real-world experiments on diverse manipulation tasks show that UniSteer adapts more efficiently than strong noise-space RL and action-space human-in-the-loop baselines, improving the success rate from 20% to 90% in 66 minutes on average across four real-world adaptation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。