用LLM控制的对手,让强化学习小模型在奇幻战斗中练战略。
Reinforcement Learning Environment with LLM-Controlled Adversary in D&D 5th Edition Combat
- 用DQN训练小智能体,对抗由GPT-4o和LLaMA 3控制的复杂对手
- 小智能体在标准指标上胜出,但LLM带来更高战略深度
- 适合研究复杂规则环境下的自适应策略与互动模拟
本研究设计并实现了一个基于《龙与地下城》5版战斗场景的强化学习(RL)环境,通过先进大语言模型(如GPT-4o和LLaMA 3 8B)控制的强对抗性代理,挑战小型RL智能体。研究采用深度Q网络(DQN)训练小型智能体,构建了一个用于战略人工智能开发的测试平台,同时具备教育价值,可模拟动态且不可预测的战斗情境。我们成功将复杂语言模型集成至强化学习框架中,提升了战略决策能力。结果显示,尽管小型智能体在标准指标上普遍优于LLM控制的对手,但LLM提供的战略深度显著增强了整体AI在这一规则密集型环境中的表现。本文探讨了该方法的创新性及其对掌握复杂环境、发展自适应策略的意义,并展望其在人工智能驱动的交互式仿真中的潜在应用。
原文摘要 · Abstract (English)
The objective of this study is to design and implement a reinforcement learning (RL) environment using D\&D 5E combat scenarios to challenge smaller RL agents through interaction with a robust adversarial agent controlled by advanced Large Language Models (LLMs) like GPT-4o and LLaMA 3 8B. This research employs Deep Q-Networks (DQN) for the smaller agents, creating a testbed for strategic AI development that also serves as an educational tool by simulating dynamic and unpredictable combat scenarios. We successfully integrated sophisticated language models into the RL framework, enhancing strategic decision-making processes. Our results indicate that while RL agents generally outperform LLM-controlled adversaries in standard metrics, the strategic depth provided by LLMs significantly enhances the overall AI capabilities in this complex, rule-based setting. The novelty of our approach and its implications for mastering intricate environments and developing adaptive strategies are discussed, alongside potential innovations in AI-driven interactive simulations. This paper aims to demonstrate how integrating LLMs can create more robust and adaptable AI systems, providing valuable insights for further research and educational applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。