让智能体通过心理认知和自我觉察持续优化策略,提升博弈胜率。
PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind
- 基于心智理论融合内外视角进行多步推理与决策
- 在动态博弈中胜过强化学习与传统代理方法
- 适合研究智能体协作、社会认知与策略演化
多智能体系统在真实世界模拟中展现出显著智能,得益于大语言模型的社会认知与知识检索能力。然而,具备推理、规划、决策与反思等认知链的智能体研究仍有限,尤其在动态交互场景中。此外,与人类不同,基于提示的响应在不确定博弈过程中难以感知心理状态并进行经验校准,易导致认知偏差。为此,我们提出PolicyEvol-Agent,一个基于大语言模型的综合性框架,能够系统性地获取他人意图,并自适应优化非理性策略以实现持续进化。该框架首先提取反思型专家模式,再融合心智理论及内部与外部视角的认知操作。仿真结果表明,其在最终博弈胜利上优于基于强化学习的模型与基于代理的方法。此外,策略演化机制在自动与人工评估中均验证了动态规则调整的有效性。
原文摘要 · Abstract (English)
Multi-agents has exhibited significant intelligence in real-word simulations with Large language models (LLMs) due to the capabilities of social cognition and knowledge retrieval. However, existing research on agents equipped with effective cognition chains including reasoning, planning, decision-making and reflecting remains limited, especially in the dynamically interactive scenarios. In addition, unlike human, prompt-based responses face challenges in psychological state perception and empirical calibration during uncertain gaming process, which can inevitably lead to cognition bias. In light of above, we introduce PolicyEvol-Agent, a comprehensive LLM-empowered framework characterized by systematically acquiring intentions of others and adaptively optimizing irrational strategies for continual enhancement. Specifically, PolicyEvol-Agent first obtains reflective expertise patterns and then integrates a range of cognitive operations with Theory of Mind alongside internal and external perspectives. Simulation results, outperforming RL-based models and agent-based methods, demonstrate the superiority of PolicyEvol-Agent for final gaming victory. Moreover, the policy evolution mechanism reveals the effectiveness of dynamic guideline adjustments in both automatic and human evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。