arXiv:2606.31958cs.RO2026-06被引 1

用语言提示优化通用机器人策略,让其学会复杂长程任务。

Adapting Generalist Robot Policies with Semantic Reinforcement Learning

论文配图:Adapting Generalist Robot Policies with Semantic Reinforcement Learning
图 1 · 摘自论文原文
  • 通过调节语言输入激活预训练技能,实现任务求解
  • 在真实与仿真环境中显著提升长程任务成功率
  • 适合需要快速适应新任务的机器人部署场景

通用机器人策略通过大规模预训练学习多样化行为。理论上,这使其成为下游任务强化学习的良好先验。然而实践中,标准强化学习方法直接优化机器人动作,要求基础策略的动作分布从一开始就接近高性能策略,这一假设在超出预训练分布的复杂或长时序任务中失效。我们的关键洞察是:对于足够表达能力强的通用策略,语言提示是解决此类任务的有效替代空间——调节语言输入可激发策略库中已有的技能,并组合完成超越零样本能力的任务。我们提出语义动作强化学习(SARL),通过在线交互学习优化该提示空间,将通用策略视为可调控的技能先验。重要的是,利用预训练技能而非从零学习新技能,实现了结构化、语义明确的探索,以及高效的在线改进;通过经验学习调节提示,使它们扎根于真实世界行为,提升任务求解鲁棒性。在真实场景和模拟基准上,SARL解锁了根本性的新能力——使视觉语言动作模型适应解决复杂长时序任务,并显著优于现有方法在部署中改进机器人行为的表现。

原文摘要 · Abstract (English)

Generalist robot policies learn a diverse repertoire of behaviors from large-scale pretraining. In principle, this makes them excellent priors for downstream adaptation via reinforcement learning (RL). In practice, however, standard RL methods leveraging this prior optimize directly over robot actions, requiring the base policy's action distribution to be close to that of a performant policy from the start. This assumption breaks down for complex or long-horizon tasks that fall outside the pretraining distribution. Our key insight is that, for sufficiently expressive generalist policies, language prompts are an effective alternative space for learning to solve such tasks: modulating language inputs elicits skills already within the policy's repertoire, which can be composed to solve tasks beyond its zero-shot capabilities. We propose Semantic Action Reinforcement Learning (SARL), which learns to optimize this prompt space through online interaction, treating the generalist policy as a controllable skill prior. Importantly, leveraging pretrained skills rather than learning new ones from scratch yields structured, semantically meaningful exploration and highly efficient online improvement, and learning to modulate prompts through experience grounds them in induced real-world behaviors for robust task-solving. Across real-world settings and simulated benchmarks, we show SARL unlocks fundamentally new capabilities -- adapting VLA behavior to solve complex, long-horizon tasks -- and significantly outperforms existing approaches for improving robot behavior in deployment.

机器人强化学习语言提示通用策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。