arXiv:2509.18463cs.ROcs.LG2025-09

通过随机扰动奖励函数,让机器人学会多样化的倒液技能。

Robotic Skill Diversification via Active Mutation of Reward Functions in Reinforcement Learning During a Liquid Pouring Task

  • 用高斯噪声扰动奖励权重,动态生成不同学习目标。
  • 策略覆盖从标准倒液到清洁、混合等新技能,行为多样性显著。
  • 适合希望机器人自主发展多任务能力的研究者。

本文研究在强化学习中主动扰动奖励函数如何促进机器人操作任务中的技能多样化,以液体倾倒为例。我们提出一种基于高斯噪声扰动奖励函数各分量权重的新框架,借鉴人类运动控制中的成本-收益权衡模型,设计包含准确性、时间与努力三项关键指标的奖励函数。实验在NVIDIA Isaac Sim仿真环境中进行,使用Franka Emika Panda机械臂持杯倾倒液体至容器。采用近端策略优化(PPO)算法,系统探索不同权重扰动配置对学习策略的影响。结果表明,所学策略表现出丰富的行为模式:既包括原定倾倒任务的变体,也涌现出适用于意外场景的新技能,如容器边缘清洁、液体混合和浇水。该方法为机器人在特定任务上实现多样化学习提供了有效路径,并可能衍生出未来任务所需的实用技能。

原文摘要 · Abstract (English)

This paper explores how deliberate mutations of reward function in reinforcement learning can produce diversified skill variations in robotic manipulation tasks, examined with a liquid pouring use case. To this end, we developed a new reward function mutation framework that is based on applying Gaussian noise to the weights of the different terms in the reward function. Inspired by the cost-benefit tradeoff model from human motor control, we designed the reward function with the following key terms: accuracy, time, and effort. The study was performed in a simulation environment created in NVIDIA Isaac Sim, and the setup included Franka Emika Panda robotic arm holding a glass with a liquid that needed to be poured into a container. The reinforcement learning algorithm was based on Proximal Policy Optimization. We systematically explored how different configurations of mutated weights in the rewards function would affect the learned policy. The resulting policies exhibit a wide range of behaviours: from variations in execution of the originally intended pouring task to novel skills useful for unexpected tasks, such as container rim cleaning, liquid mixing, and watering. This approach offers promising directions for robotic systems to perform diversified learning of specific tasks, while also potentially deriving meaningful skills for future tasks.

机器人强化学习技能多样化奖励函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。