用大模型自动设计游戏生成奖励,减少人工干预。
PCGRLLM: Large Language Model-Driven Reward Design for Procedural Content Generation Reinforcement Learning
- 基于大模型和反馈机制生成奖励函数,支持程序化内容生成
- 在二维环境中实现接近人类水平的奖励生成性能
- 适合游戏开发、AI创作等需快速生成奖励的任务
奖励设计在游戏AI训练中至关重要,但通常需要大量领域知识和人力投入。近年来,已有研究探索利用大语言模型(LLMs)生成奖励以训练游戏智能体或控制机器人。在内容生成领域,早期工作已尝试为强化学习智能体生成奖励函数。本文提出PCGRLLM,一种基于前期工作的扩展架构,采用反馈机制与多种基于推理的提示工程方法。我们在一个二维环境中的故事到奖励生成任务上,使用两个先进的LLM,评估了多种推理型提示方法。实验结果揭示了LLMs在内容生成任务中的关键能力,相较于先前结构有显著性能提升,达到接近人类的水平。本工作展示了降低游戏AI开发中人工依赖的潜力,同时支持并增强创造性过程。
原文摘要 · Abstract (English)
Reward design plays a pivotal role in the training of game AIs, requiring substantial domain-specific knowledge and human effort. In recent years, several studies have explored reward generation for training game agents and controlling robots using large language models (LLMs). In the content generation literature, there has been early work on generating reward functions for reinforcement learning agent generators. This work introduces PCGRLLM, an extended architecture based on earlier work, which employs a feedback mechanism and several reasoning-based prompt engineering techniques. We evaluate the proposed method on a story-to-reward generation task in a two-dimensional environment using two state-of-the-art LLMs across various reasoning-based prompting methods. Our experiments provide insightful evaluations that demonstrate the capabilities of LLMs essential for content generation tasks. The results demonstrate a substantial performance improvement over the previous structure, achieving performance comparable to that of humans. Our work demonstrates the potential to reduce human dependency in game AI development, while supporting and enhancing creative processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。