让智能体按优先级分层满足需求,提升整体奖励表现。
Creating Hierarchical Dispositions of Needs in an Agent
- 用多输出奖励代理构建分层目标优先级
- 在摆锤环境中达到当前最优性能
- 适合研究目标优先级与决策层次的学者
我们提出一种学习分层抽象的新方法,用于对竞争性目标进行优先排序,从而提升全局期望奖励。该方法采用一个具有多个标量输出的辅助奖励智能体,每个输出对应一个抽象层级。主智能体则以分层方式学习最大化这些输出,并使每一层依赖于前一层的最优实现。我们推导出一个方程,按优先级排序各标量值与全局奖励,形成指导目标生成的需要层次结构。在Pendulum v1环境中的实验表明,该方法性能优于基线实现,并达到了当前最优结果。
原文摘要 · Abstract (English)
We present a novel method for learning hierarchical abstractions that prioritize competing objectives, leading to improved global expected rewards. Our approach employs a secondary rewarding agent with multiple scalar outputs, each associated with a distinct level of abstraction. The traditional agent then learns to maximize these outputs in a hierarchical manner, conditioning each level on the maximization of the preceding level. We derive an equation that orders these scalar values and the global reward by priority, inducing a hierarchy of needs that informs goal formation. Experimental results on the Pendulum v1 environment demonstrate superior performance compared to a baseline implementation.We achieved state of the art results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。