用内在价值驱动强化学习,让智能体像人一样有自主需求和目标。
Innate-Values-driven Reinforcement Learning based Cognitive Modeling
- 基于内在价值与预期效用理论构建新强化学习模型
- 在VIZDoom平台比传统算法性能更优,能合理分配多种需求
- 适合研究具身智能、个性化决策与长期社会协作的场景
内在价值描述了智能体的内在动机,反映其对目标追求的固有偏好与兴趣,驱动其发展多样技能以满足不同需求。传统强化学习依赖环境反馈奖励,但现实中奖励由智能体自身的内在价值系统生成,个体间差异显著。将智能体视为自组织系统,通过平衡内外部效用、根据任务需求提升自我认知,是实现长期支持他人、融入群体并保障安全和谐的关键。为此,我们提出一种基于联合动机模型与预期效用理论的内在价值驱动强化学习(IVRL)模型,用于模拟智能体在决策与学习中行为演化的复杂性。进一步设计了基于IVRL的IV-DQN与IV-A2C两个模型。在角色扮演游戏测试平台VIZDoom上,与DQN、DDQN、A2C及PPO等基准算法对比,结果表明,基于IVRL的模型能有效帮助智能体理性组织多重需求,实现更优性能。
原文摘要 · Abstract (English)
Innate values describe agents' intrinsic motivations, which reflect their inherent interests and preferences for pursuing goals and drive them to develop diverse skills that satisfy their various needs. Traditional reinforcement learning (RL) is learning from interaction based on the feedback rewards of the environment. However, in real scenarios, the rewards are generated by agents' innate value systems, which differ vastly from individuals based on their needs and requirements. In other words, considering the AI agent as a self-organizing system, developing its awareness through balancing internal and external utilities based on its needs in different tasks is a crucial problem for individuals learning to support others and integrate community with safety and harmony in the long term. To address this gap, we propose a new RL model termed innate-values-driven RL (IVRL) based on combined motivations' models and expected utility theory to mimic its complex behaviors in the evolution through decision-making and learning. Then, we introduce two IVRL-based models: IV-DQN and IV-A2C. By comparing them with benchmark algorithms such as DQN, DDQN, A2C, and PPO in the Role-Playing Game (RPG) reinforcement learning test platform VIZDoom, we demonstrated that the IVRL-based models can help the agent rationally organize various needs, achieve better performance effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。